PRISM

Population-Retrieved Inference for Structured Multimodal Prognosis

1Eindhoven University of Technology   2University of Melbourne
Published in Transactions on Machine Learning Research (TMLR), 10/2026
Overview of the PRISM architecture. Five modality streams — structured EHR, current ECG, longitudinal ECG history, retrieved EHR context, and retrieved ECG context — are independently encoded and projected as soft tokens into a shared sequence consumed by a LoRA-adapted MedGemma-4B decoder. The population memory bank retrieves the k=3 nearest training-fold encounters from dual EHR and ECG indices. The temporal history aggregator pools prior ECG embeddings under exponential recency weighting.
Fig. 1 — Five modality streams (structured EHR, current ECG, longitudinal ECG history, retrieved EHR context, retrieved ECG context) are independently encoded and projected as soft tokens into a shared sequence consumed by a LoRA-adapted MedGemma-4B decoder, with a dual-modality population memory bank and a recency-weighted history aggregator.
The big idea

Sensor-language models stop at perception. Prognosis needs more evidence than that.

Clinical prognosis requires integrating heterogeneous evidence from structured measurements, physiological waveforms, longitudinal history, and population context — in one model, without discarding any of it.

01Numbers, kept as numbers

Multimodal clinical LLMs typically serialize labs, vitals, and demographics into text strings, so a troponin of 0.11 and a troponin of 0.01 become subword sequences whose clinical distance is opaque to the embedding layer. PRISM instead runs all 470 structured clinical features through an FT-Transformer, mapping every feature to its own learned token and preserving the ordinal, ratio-scale structure that text serialization discards.

02The record has a history

Existing multimodal systems encode only the index-visit ECG, treating every encounter as the patient’s first. PRISM pools prior-visit ECG embeddings under exponential recency weighting — a single learned decay scalar — injecting serial waveform change into the model at negligible parameter cost.

03A memory of the population

Even large language models reason only from what they memorized in pre-training, with no way to consult clinically similar patients. PRISM indexes per-encounter ECG and EHR embeddings from the training cohort in a dual-modality memory bank, and supplies the mean-pooled representations of the k=3 nearest neighbors as additional soft-token context — population evidence in representation space, with no individual record ever entering the prompt.

“Can structured records, longitudinal waveforms, and population context be integrated through a single tokenized interface, with no task-specific heads?”

How it works

Five streams, one shared sequence, no task-specific heads

A single trained model answers all 1,443 task questions, each in its own forward pass.

  1. 01

    Encode every stream

    An FT-Transformer tokenizes all 470 structured EHR features individually; a frozen D-BETA ECG encoder processes the current visit and the longitudinal history.

  2. 02

    Aggregate the ECG history

    Prior-visit embeddings are pooled under exponential recency weighting (wi ∝ exp(−|λ|Δt)), compressing the 96-hour pre-arrival window into one vector.

  3. 03

    Retrieve from the population

    Dual EHR and ECG memory banks return the k=3 nearest training-fold encounters; their mean-pooled embeddings — not labels — join the sequence as soft tokens.

  4. 04

    Reason and answer

    A LoRA-adapted MedGemma-4B decoder consumes all tokens and answers via a generative Yes/No pathway or a classification head — one forward pass per question, no task-specific output heads.

What we found

Strongest reported AUROC across every MDS-ED task group

Evaluated on the 1,443-task MDS-ED benchmark (MIMIC-IV), against 15 baselines spanning five model families.

0
Diagnosis AUROC (%)
1,428 ICD-10 tasks
0
Deterioration AUROC (%)
6 tasks
0
ICU admission AUROC (%)
2 tasks
0
Mortality AUROC (%)
7 tasks
PRISM performance vs. MDS-ED baseline across clinical prediction categories: mortality at multiple horizons, clinical deterioration endpoints, ICU admission, and ICD-chapter-level diagnosis groups. Gains are largest on laboratory-dependent outcomes such as blood and immune disorders and musculoskeletal conditions, and on deterioration endpoints where metabolic measurements are clinically relevant.
Fig. 2 — PRISM vs. MDS-ED baseline. Gains are largest on laboratory-dependent outcomes (blood & immune disorders +6.06%, musculoskeletal +6.33%) and on deterioration endpoints where metabolic measurements are clinically relevant (severe hypoxemia +13.05%).
Table 1 — Benchmark results across all 1,443 MDS-ED tasks. AUROC (%), macro-averaged within each task group.
MethodDiagnosisDeteriorationICUMortality
ECG foundation models · waveform input
HeartBERT56.3558.2457.2861.11
ECGFM-KED58.7564.7661.3265.76
ESI60.9863.3462.5266.28
ECGFounder61.5269.3564.1369.42
Image-based models · rendered ECG input
GEM61.5255.8060.4065.70
PULSE58.4761.1064.7068.70
TimeMaster57.9160.8063.8057.30
ECG-Instruct47.3346.0044.4043.70
Signal-based QA models · ECG waveform input
ECGQA-FSL50.5650.4651.0652.15
Q-HEART50.1555.4254.8956.32
ECG-Chat49.7057.5456.4055.78
Frontier LLMs · zero-shot, text-only input
GPT-555.6878.1270.0462.27
Gemini-2.5 Pro57.8278.7270.7072.38
Multimodal baselines · waveform + EHR
MDS-ED82.5690.7090.6391.68
UniPACT83.9891.1790.5091.82
PRISM (ours) · waveform + EHR + history + retrieval
PRISMgen86.2492.9091.7992.82
PRISMdisc86.0693.1291.5592.79

Single-modality ECG models and zero-shot frontier LLMs operate well below the multimodal tier, confirming that waveform or parametric knowledge alone cannot substitute for task-specific multimodal supervision. PRISM surpasses MDS-ED and UniPACT on every task group, by 2.26 points on Diagnosis, 1.95 on Deterioration, 1.29 on ICU admission, and 1.00 on Mortality. Counting tasks whose lower 95% confidence bound exceeds 0.80 AUROC, PRISM qualifies on 1,145 of 1,443 tasks (79.3%), against 883 for UniPACT (61.2%) and 623 for MDS-ED (43.2%).

External validity

Transfers to a new hospital, then adapts with supervision

MC-MED — Stanford emergency-department visits (2020–2022), a separate institution from MIMIC-IV, with a single-lead ECG repeated across the 12 encoder channels.

Table 2 — External evaluation on MC-MED. All entries are AUROC; PRISM uses the generative pathway.
Model / conditionDiagnosisDeteriorationICUMortalityOverall AUROC
Source-only transfer
PRISM, original bank0.70860.71120.81880.89770.7841 [0.751, 0.814]
Local-bank adaptation · no parameter updates
PRISM, MC-MED bank0.70800.71330.81930.90250.7858 [0.754, 0.815]
Supervised local training / adaptation
Gradient boosting0.70120.75090.89240.94650.8227 [0.796, 0.851]
Fusion MLP0.61840.62270.71880.81500.6937 [0.658, 0.729]
Fusion MLP + retrieval0.58470.64710.69870.74690.6694 [0.635, 0.703]
Fine-tuned PRISM0.84450.76900.88300.94790.8611 [0.834, 0.887]
Inference-time controls of fine-tuned PRISM
Retrieval zeroed0.79200.73890.86810.92150.8301 [0.802, 0.857]
Random retrieval0.79190.73100.86570.92360.8280 [0.799, 0.856]

Supervised fine-tuning reaches 0.8611 overall AUROC, +7.70 points over source-only transfer, and exceeds locally trained gradient boosting by 3.84 points overall. Zeroing retrieval lowers fine-tuned AUROC by 3.10 points and random neighbors by 3.31 — genuine retrieval has higher point estimates in all four task groups.

What the memory bank actually retrieves

A test encounter and the three training-fold neighbors nearest it in the dual-modality index.

Qualitative analysis of the population memory bank. A test encounter with ICD-10 label Z79 (long-term drug therapy, positive outcome) and three training-fold encounters retrieved from the dual-modality memory bank, each showing an ECG strip and key clinical features with abnormality flags.
Fig. 3 — Qualitative analysis of the population memory bank.

Test patient · ICD-10 Z79, positive outcome

Elevated heart rate (101 bpm) and elevated glucose (114 mg/dL) are flagged abnormal alongside normal SpO₂, respiratory rate, and temperature.

Three nearest neighbors, independently indexed

The EHR and ECG indices are queried independently; all three displayed encounters are returned by both and share the same discharge label and positive outcome. Index similarity: 96.5%, 76.7%, and 99.5%.

What the model actually receives

The decoder consumes the two mean-pooled retrieval vectors defined in the paper’s Equation 5, not these individual records — so provenance can be reviewed offline without ever placing a retrieved patient’s data in the prompt. Shared abnormalities (elevated glucose, elevated heart rate) illustrate local coherence in the learned embedding space, not cohort-level inference.

The benchmark

MDS-ED: 1,443 tasks over raw 12-lead ECG and structured EHR

Evaluated on a held-out MIMIC-IV test set of 6,454 patient records, following Stage 1 and Stage 2 optimization on 4.5 million training samples.

0
benchmark tasks
0
structured EHR features
0
held-out test patients
0
training samples
Dx

Diagnosis

1,428 ICD-10 discharge tasks, predicted at the end of the hospital stay.

86.24% AUROC
Dt

Deterioration

6 tasks predicting clinical deterioration from data available at ED arrival.

93.12% AUROC
ICU

ICU admission

2 tasks: 24-hour and overall ICU admission.

91.79% AUROC
Mx

Mortality

7 tasks spanning in-hospital and multi-horizon mortality.

92.82% AUROC
The paper

Heterogeneous evidence, one tokenized interface

Full paper, code, and model weights are on their way — links will be added here once available.

BibTeX
@article{shafiq2026prism,
  title   = {Retrieval-Augmented Multimodal Language Model with
             Heterogeneous Evidence Integration for Clinical Prognosis},
  author  = {Shafiq, Hamza and Tang, Jialu and Dang, Ting and Hu, Jun and Saeed, Aaqib},
  journal = {Transactions on Machine Learning Research},
  year    = {2026}
}