Diagnosis
1,428 ICD-10 discharge tasks, predicted at the end of the hospital stay.
86.24% AUROCPopulation-Retrieved Inference for Structured Multimodal Prognosis
1Eindhoven University of Technology 2University of Melbourne
Clinical prognosis requires integrating heterogeneous evidence from structured measurements, physiological waveforms, longitudinal history, and population context — in one model, without discarding any of it.
Multimodal clinical LLMs typically serialize labs, vitals, and demographics into text strings, so a troponin of 0.11 and a troponin of 0.01 become subword sequences whose clinical distance is opaque to the embedding layer. PRISM instead runs all 470 structured clinical features through an FT-Transformer, mapping every feature to its own learned token and preserving the ordinal, ratio-scale structure that text serialization discards.
Existing multimodal systems encode only the index-visit ECG, treating every encounter as the patient’s first. PRISM pools prior-visit ECG embeddings under exponential recency weighting — a single learned decay scalar — injecting serial waveform change into the model at negligible parameter cost.
Even large language models reason only from what they memorized in pre-training, with no way to consult clinically similar patients. PRISM indexes per-encounter ECG and EHR embeddings from the training cohort in a dual-modality memory bank, and supplies the mean-pooled representations of the k=3 nearest neighbors as additional soft-token context — population evidence in representation space, with no individual record ever entering the prompt.
“Can structured records, longitudinal waveforms, and population context be integrated through a single tokenized interface, with no task-specific heads?”
A single trained model answers all 1,443 task questions, each in its own forward pass.
An FT-Transformer tokenizes all 470 structured EHR features individually; a frozen D-BETA ECG encoder processes the current visit and the longitudinal history.
Prior-visit embeddings are pooled under exponential recency weighting (wi ∝ exp(−|λ|Δt)), compressing the 96-hour pre-arrival window into one vector.
Dual EHR and ECG memory banks return the k=3 nearest training-fold encounters; their mean-pooled embeddings — not labels — join the sequence as soft tokens.
A LoRA-adapted MedGemma-4B decoder consumes all tokens and answers via a generative Yes/No pathway or a classification head — one forward pass per question, no task-specific output heads.
Evaluated on the 1,443-task MDS-ED benchmark (MIMIC-IV), against 15 baselines spanning five model families.
| Method | Diagnosis | Deterioration | ICU | Mortality |
|---|---|---|---|---|
| ECG foundation models · waveform input | ||||
| HeartBERT | 56.35 | 58.24 | 57.28 | 61.11 |
| ECGFM-KED | 58.75 | 64.76 | 61.32 | 65.76 |
| ESI | 60.98 | 63.34 | 62.52 | 66.28 |
| ECGFounder | 61.52 | 69.35 | 64.13 | 69.42 |
| Image-based models · rendered ECG input | ||||
| GEM | 61.52 | 55.80 | 60.40 | 65.70 |
| PULSE | 58.47 | 61.10 | 64.70 | 68.70 |
| TimeMaster | 57.91 | 60.80 | 63.80 | 57.30 |
| ECG-Instruct | 47.33 | 46.00 | 44.40 | 43.70 |
| Signal-based QA models · ECG waveform input | ||||
| ECGQA-FSL | 50.56 | 50.46 | 51.06 | 52.15 |
| Q-HEART | 50.15 | 55.42 | 54.89 | 56.32 |
| ECG-Chat | 49.70 | 57.54 | 56.40 | 55.78 |
| Frontier LLMs · zero-shot, text-only input | ||||
| GPT-5 | 55.68 | 78.12 | 70.04 | 62.27 |
| Gemini-2.5 Pro | 57.82 | 78.72 | 70.70 | 72.38 |
| Multimodal baselines · waveform + EHR | ||||
| MDS-ED | 82.56 | 90.70 | 90.63 | 91.68 |
| UniPACT | 83.98 | 91.17 | 90.50 | 91.82 |
| PRISM (ours) · waveform + EHR + history + retrieval | ||||
| PRISMgen | 86.24 | 92.90 | 91.79 | 92.82 |
| PRISMdisc | 86.06 | 93.12 | 91.55 | 92.79 |
Single-modality ECG models and zero-shot frontier LLMs operate well below the multimodal tier, confirming that waveform or parametric knowledge alone cannot substitute for task-specific multimodal supervision. PRISM surpasses MDS-ED and UniPACT on every task group, by 2.26 points on Diagnosis, 1.95 on Deterioration, 1.29 on ICU admission, and 1.00 on Mortality. Counting tasks whose lower 95% confidence bound exceeds 0.80 AUROC, PRISM qualifies on 1,145 of 1,443 tasks (79.3%), against 883 for UniPACT (61.2%) and 623 for MDS-ED (43.2%).
MC-MED — Stanford emergency-department visits (2020–2022), a separate institution from MIMIC-IV, with a single-lead ECG repeated across the 12 encoder channels.
| Model / condition | Diagnosis | Deterioration | ICU | Mortality | Overall AUROC |
|---|---|---|---|---|---|
| Source-only transfer | |||||
| PRISM, original bank | 0.7086 | 0.7112 | 0.8188 | 0.8977 | 0.7841 [0.751, 0.814] |
| Local-bank adaptation · no parameter updates | |||||
| PRISM, MC-MED bank | 0.7080 | 0.7133 | 0.8193 | 0.9025 | 0.7858 [0.754, 0.815] |
| Supervised local training / adaptation | |||||
| Gradient boosting | 0.7012 | 0.7509 | 0.8924 | 0.9465 | 0.8227 [0.796, 0.851] |
| Fusion MLP | 0.6184 | 0.6227 | 0.7188 | 0.8150 | 0.6937 [0.658, 0.729] |
| Fusion MLP + retrieval | 0.5847 | 0.6471 | 0.6987 | 0.7469 | 0.6694 [0.635, 0.703] |
| Fine-tuned PRISM | 0.8445 | 0.7690 | 0.8830 | 0.9479 | 0.8611 [0.834, 0.887] |
| Inference-time controls of fine-tuned PRISM | |||||
| Retrieval zeroed | 0.7920 | 0.7389 | 0.8681 | 0.9215 | 0.8301 [0.802, 0.857] |
| Random retrieval | 0.7919 | 0.7310 | 0.8657 | 0.9236 | 0.8280 [0.799, 0.856] |
Supervised fine-tuning reaches 0.8611 overall AUROC, +7.70 points over source-only transfer, and exceeds locally trained gradient boosting by 3.84 points overall. Zeroing retrieval lowers fine-tuned AUROC by 3.10 points and random neighbors by 3.31 — genuine retrieval has higher point estimates in all four task groups.
A test encounter and the three training-fold neighbors nearest it in the dual-modality index.
Elevated heart rate (101 bpm) and elevated glucose (114 mg/dL) are flagged abnormal alongside normal SpO₂, respiratory rate, and temperature.
The EHR and ECG indices are queried independently; all three displayed encounters are returned by both and share the same discharge label and positive outcome. Index similarity: 96.5%, 76.7%, and 99.5%.
The decoder consumes the two mean-pooled retrieval vectors defined in the paper’s Equation 5, not these individual records — so provenance can be reviewed offline without ever placing a retrieved patient’s data in the prompt. Shared abnormalities (elevated glucose, elevated heart rate) illustrate local coherence in the learned embedding space, not cohort-level inference.
Evaluated on a held-out MIMIC-IV test set of 6,454 patient records, following Stage 1 and Stage 2 optimization on 4.5 million training samples.
1,428 ICD-10 discharge tasks, predicted at the end of the hospital stay.
86.24% AUROC6 tasks predicting clinical deterioration from data available at ED arrival.
93.12% AUROC2 tasks: 24-hour and overall ICU admission.
91.79% AUROC7 tasks spanning in-hospital and multi-horizon mortality.
92.82% AUROCFull paper, code, and model weights are on their way — links will be added here once available.
@article{shafiq2026prism,
title = {Retrieval-Augmented Multimodal Language Model with
Heterogeneous Evidence Integration for Clinical Prognosis},
author = {Shafiq, Hamza and Tang, Jialu and Dang, Ting and Hu, Jun and Saeed, Aaqib},
journal = {Transactions on Machine Learning Research},
year = {2026}
}