arXiv:2609.20827v1 Announce Type: new Abstract: Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.
From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
Researchers introduced DischargeBench, a persona-grounded simulation that evaluates large language models as hospital discharge educators through multi-turn dialogue with a Virtual Patient, with an Education Monitor Agent regulating patient realism without modifying the educator. The team curated MIMIC-IV-Ext-DischargeBench, comprising 477 cases across 24 ICD chapters with persona axes covering personality, education level, health literacy, and past-medical-history recall, scoring each simulation on Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency via an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores concealed clinically relevant variation across ICD chapters and patient personas, with difficult personas exposing coverage failures, comprehension gaps, and reduced source-answer agreement, leading the authors to argue that LLM evaluation for discharge education should center patient understanding rather than text quality or answer accuracy alone.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.