cd /news/artificial-intelligence/large-language-models-exhibit-unreli… · home › topics › artificial-intelligence › article
[ARTICLE · art-145219] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

A study of matched intensive-care trajectories from electronic health records found that large language models update clinical judgments unreliably as patient evidence evolves, with conditioning on a preceding judgment more often increasing than reducing prediction error when estimates changed. The arXiv paper (2610.02684v1) reports two failure modes: models responded more strongly to worsening than to matched improving respiratory evidence, and raising prior risk from 10% to 90% shifted estimates by 26.2 percentage points, while prompting did not restore reliable updating. The authors propose Evidence-Validated Longitudinal Update (EVLU), which identified fewer but more reliable revisions, exposing a reliability-coverage trade-off.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.02684v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-model…] indexed:0 read:1min 2026-10-05 · —