{"slug": "mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm", "title": "MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation", "summary": "Researchers released MTDiag, a multi-turn diagnostic dialogue dataset built from DDXPlus, MIMIC-IV, and published case reports, to evaluate large language models (LLMs) as diagnostic agents in dynamic clinical encounters. The dataset, normalized to UMLS concept identifiers and ICD-10 codes, includes a UserLM-8B-based utterance-generation pipeline and physician-validated cases, alongside new clinical knowledge-grounded metrics beyond diagnostic accuracy.", "body_md": "arXiv:2608.25085v1 Announce Type: new\nAbstract: Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.", "url": "https://wpnews.pro/news/mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm", "canonical_source": "https://arxiv.org/abs/2608.25085", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:20:08.666674+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["MTDiag", "DDXPlus", "MIMIC-IV", "AJCR", "UMLS", "ICD-10", "UserLM-8B"], "alternates": {"html": "https://wpnews.pro/news/mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm", "markdown": "https://wpnews.pro/news/mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm.md", "text": "https://wpnews.pro/news/mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm.txt", "jsonld": "https://wpnews.pro/news/mtdiag-a-multi-turn-diagnostic-dataset-towards-clinically-meaningful-llm.jsonld"}}