cd /news/large-language-models/llms-struggle-to-simulate-human-beli… · home topics large-language-models article
[ARTICLE · art-81448] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LLMs struggle to simulate human belief updates in controlled environments

A new study from arXiv (2607.28347v1) found that six large language models, including Qwen3-32B and GPT-5-Mini, fail to simulate human belief updates in controlled environments, with only some matching post-stance distributions when given participants' actual initial stances. The study, which compared LLM outputs against data from 391 UK participants on Prolific, identified three systematic biases: overrepresentation of neutral positions, more frequent but smaller belief shifts, and failure to rank comments by convincingness.

read1 min views1 publishedJul 31, 2026

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llms-struggle-to-sim…] indexed:0 read:1min 2026-07-31 ·