An hourly sepsis score learned from routinely charted intensive-care data separated patients who survived from those who did not across four starting levels of illness severity. On its fixed 0-to-10 scale, the average gap between non-survivors and survivors ranged from 1.19 to 1.64 points in the four groups.
The score was learned from pairwise outcome rankings rather than from a treatment intervention. It does not show that using the score would improve care, because no score-guided intervention was tested.
A score learned from the record
The study included 29,116 adults in MIMIC-IV and 7,691 in Emory Healthcare, all meeting Sepsis-3 criteria. It used 43 routinely charted physiologic, laboratory and treatment variables over a 72-hour treatment window, with no intervention administered for the study.
Researchers trained a multilayer perceptron, a type of neural network, using pairwise rankings of patient stays. In the mortality ranking, stays ending in survival were preferred to stays ending in death. They also tested rankings that added treatment intensity or matched patients by baseline severity.
Within each hospital system, 20% of patients were kept as a permanent test set. That holdout contained 5,823 MIMIC-IV admissions and 1,538 Emory admissions; the remaining data were divided into four mortality-stratified validation folds. Observations were aggregated by hour, restricted to plausible ranges, carried forward within an admission when missing, and any remaining gaps were filled using the median from the training portion.
The signal followed some physiological changes
Among 1,854 held-out MIMIC-IV patients in the lactate analysis, score change tracked lactate change. The Spearman rank correlation, a measure of how two changes move together in order, was 0.39, and the stratified correlation rose from 0.25 to 0.55 as baseline lactate severity increased.
The score's changes also moved inversely with changes in mean arterial pressure, or MAP. This relationship was strongest below 65 mmHg and weakest above 85 mmHg; MAP and creatinine associations were much weaker than the lactate association.
The mortality ranking was chosen as the final model because it used the simplest ranking approach, produced the strongest agreement across repeated model fits, and showed similar relationships with established severity indices.
Transfer between hospitals was the harder test
Score behavior agreed more within the same hospital system than across institutions. Median patient-level correlations between the two sites were 0.54 for MIMIC-IV patients, with a 95% confidence interval from 0.53 to 0.55, and 0.59 for Emory patients, with an interval from 0.56 to 0.61. The within-site ceilings were 0.92 and 0.90, while negative correlations appeared for 14.2% and 12.7% of patients, respectively.
Mean correlations with established severity indices ranged from 0.25 to 0.46 in MIMIC-IV and from 0.18 to 0.59 at Emory. Across 24 ablation conditions at each site, no mean correlation exceeded 0.08 in absolute value.
For mortality discrimination, the score's median per-model area under the receiver operating characteristic curve, or AUROC, was 0.764 in MIMIC-IV and 0.742 at Emory. The 20-model consensus was close to a pointwise mortality-trained neural-network baseline: 0.791 versus 0.789 in MIMIC-IV, and 0.765 versus 0.768 at Emory.
For mortality ranking, random-seed agreement was 0.79 at Emory and 0.78 in MIMIC-IV. Cross-institution agreement was 0.63 and 0.56, retaining 77% and 70% of within-site agreement. Adding treatment intensity reduced seed agreement to 0.51 and 0.67 and cross-site agreement to 0.44 and 0.38.
A research signal, not a bedside score yet
Because no intervention was administered, the study did not test whether score-guided decisions change treatment or outcomes. Uncertainty was estimated with whole-patient bootstrap resampling, using 95% percentile intervals based on either 1,000 or 200 replicates. Spearman correlations were averaged in Fisher z space.
The 43-input state omitted organ-support variables, making it harder to distinguish illness severity from the effects of treatment. The creatinine-score correlation was 0.19 in the baseline 1.2 to 2.0 mg/dL range but fell to 0.06 above 2.0 mg/dL, with an interval spanning zero.
The lower cross-institution agreement leaves the score's behavior outside these two hospital systems unresolved. The study did not test score-guided clinical decisions or patient outcomes, so its results describe associations in observed care rather than the effects of using the score.
The findings point to the need for local validation and prospective testing before the score can be interpreted at the bedside. For now, it is a research measure of ICU trajectories whose transfer between hospitals remains imperfect.
Paper data and sources
Original title: Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
Authors: Kevin Zhu, Ryan Zhang, Baraa Abed et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-27
DOI: Not available
Original paper · Full text