cd /news/artificial-intelligence/earlydx-an-admission-anchored-benchm… · home topics artificial-intelligence article
[ARTICLE · art-84225] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

Researchers introduced EarlyDx, a benchmark built from 154,834 emergency department encounters in MIMIC-IV, to evaluate open-ended early diagnosis at hospital admission. The study found that no evaluated system, including frontier general, medical-specialized, or in-domain post-trained models, reliably synthesizes admission-time evidence, with zero-shot models recovering only 3-31% of diagnoses requiring inference, and post-training raising inference-dependent recall to 56% but still leaving a significant gap.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @earlydx 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/earlydx-an-admission…] indexed:0 read:1min 2026-08-03 ·