{"slug": "earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed", "title": "EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses", "summary": "Researchers introduced EarlyDx, a benchmark built from 154,834 emergency department encounters in MIMIC-IV, to evaluate open-ended early diagnosis at hospital admission. The study found that no evaluated system, including frontier general, medical-specialized, or in-domain post-trained models, reliably synthesizes admission-time evidence, with zero-shot models recovering only 3-31% of diagnoses requiring inference, and post-training raising inference-dependent recall to 56% but still leaving a significant gap.", "body_md": "arXiv:2607.28788v1 Announce Type: new\nAbstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.", "url": "https://wpnews.pro/news/earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed", "canonical_source": "https://arxiv.org/abs/2607.28788", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:13:28.478895+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["EarlyDx", "MIMIC-IV"], "alternates": {"html": "https://wpnews.pro/news/earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed", "markdown": "https://wpnews.pro/news/earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed.md", "text": "https://wpnews.pro/news/earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed.txt", "jsonld": "https://wpnews.pro/news/earlydx-an-admission-anchored-benchmark-for-open-ended-generation-of-evidence-ed.jsonld"}}