{"slug": "from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as", "title": "From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators", "summary": "Researchers introduced DischargeBench, a persona-grounded simulation that evaluates large language models as hospital discharge educators through multi-turn dialogue with a Virtual Patient, with an Education Monitor Agent regulating patient realism without modifying the educator. The team curated MIMIC-IV-Ext-DischargeBench, comprising 477 cases across 24 ICD chapters with persona axes covering personality, education level, health literacy, and past-medical-history recall, scoring each simulation on Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency via an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores concealed clinically relevant variation across ICD chapters and patient personas, with difficult personas exposing coverage failures, comprehension gaps, and reduced source-answer agreement, leading the authors to argue that LLM evaluation for discharge education should center patient understanding rather than text quality or answer accuracy alone.", "body_md": "arXiv:2609.20827v1 Announce Type: new \nAbstract: Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.", "url": "https://wpnews.pro/news/from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as", "canonical_source": "https://arxiv.org/abs/2609.20827", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 04:23:24.007090+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety", "natural-language-processing"], "entities": ["DischargeBench", "MIMIC-IV-Ext-DischargeBench", "Education Monitor Agent", "Virtual Patient", "ICD"], "alternates": {"html": "https://wpnews.pro/news/from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as", "markdown": "https://wpnews.pro/news/from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as.md", "text": "https://wpnews.pro/news/from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as.txt", "jsonld": "https://wpnews.pro/news/from-discharge-notes-to-patient-understanding-persona-grounded-open-ended-of-as.jsonld"}}