cd /news/large-language-models/llms-are-not-consistently-bayesian-q… · home topics large-language-models article
[ARTICLE · art-114360] src=machinelearning.apple.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

Researchers from Stanford University and Apple found that large language models (LLMs) are not consistently Bayesian when updating probabilistic beliefs from evidence, with non-Bayesian heuristic updates often outperforming exact Bayesian updates in downstream task performance. The study introduces a technique to quantify the information processing gap—the deviation from Bayes updates—and suggests it can serve as a diagnostic for LLM-powered inferential systems.

read1 min views1 publishedAug 28, 2026
LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs
Image: Apple ML Research

Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decisions. We introduce the novel technique of studying LLMs as information processing rules and utilize the information processing gap—the deviation from Bayes updates—to study the internal (in)consistencies of how LLMs update their probabilistic beliefs from evidence. Our extensive experiments evaluate multiple approaches in which LLMs can incorporate evidence into their beliefs. Some of these approaches produce (nearly) Bayesian updates, thus optimally processing evidence; others use a learned heuristic. Surprisingly, the non-Bayesian heuristic updates often outperform exact Bayesian updates (optimal information processing) in terms of downstream task performance—indicating the LLMs’ probabilistic models of the world are misspecified. Lastly, we show how our measure can provide diagnostics to identify issues with LLM-powered inferential systems.

  • † Stanford University
  • ‡ Equal contribution
  • ** Work done while at Apple
── more in #large-language-models 4 stories · sorted by recency
── more on @stanford university 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llms-are-not-consist…] indexed:0 read:1min 2026-08-28 ·