{"slug": "medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large", "title": "MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models", "summary": "Researchers released MedLoCoMo, a medical long-context benchmark for evaluating large language models on patient-specific clinical reasoning over multi-admission dialogues, built from deidentified MIMIC-IV records. The benchmark contains 100 patient timelines averaging 1,669.8 turns and 74,512.2 tokens per conversation, and tests single-admission, cross-admission, and adversarial unanswerable settings. Cross-admission reasoning proved consistently harder than localized evidence use across all evaluated baselines.", "body_md": "arXiv:2607.22566v1 Announce Type: new\nAbstract: MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods. The code and MedLoCoMo benchmark release is available at https://github.com/leozzy13/MedLoCoMo for use and reproducibility.", "url": "https://wpnews.pro/news/medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large", "canonical_source": "https://arxiv.org/abs/2607.22566", "published_at": "2026-07-28 04:00:00+00:00", "updated_at": "2026-07-28 04:29:23.783127+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing"], "entities": ["MedLoCoMo", "MIMIC-IV", "MIMIC-IV-Note", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large", "markdown": "https://wpnews.pro/news/medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large.md", "text": "https://wpnews.pro/news/medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large.txt", "jsonld": "https://wpnews.pro/news/medlocomo-a-long-context-multi-session-medical-dialogue-benchmark-for-large.jsonld"}}