{"slug": "large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as", "title": "Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves", "summary": "A study of matched intensive-care trajectories from electronic health records found that large language models update clinical judgments unreliably as patient evidence evolves, with conditioning on a preceding judgment more often increasing than reducing prediction error when estimates changed. The arXiv paper (2610.02684v1) reports two failure modes: models responded more strongly to worsening than to matched improving respiratory evidence, and raising prior risk from 10% to 90% shifted estimates by 26.2 percentage points, while prompting did not restore reliable updating. The authors propose Evidence-Validated Longitudinal Update (EVLU), which identified fewer but more reliable revisions, exposing a reliability-coverage trade-off.", "body_md": "arXiv:2610.02684v1 Announce Type: cross \nAbstract: Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.", "url": "https://wpnews.pro/news/large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as", "canonical_source": "https://www.machinebrief.com/news/large-language-models-exhibit-unreliable-updating-of-clinica-2uvq", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 06:12:30.710597+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "Evidence-Validated Longitudinal Update", "EVLU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as", "markdown": "https://wpnews.pro/news/large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as.md", "text": "https://wpnews.pro/news/large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as.txt", "jsonld": "https://wpnews.pro/news/large-language-models-exhibit-unreliable-updating-of-clinical-judgment-as.jsonld"}}