cd /news/large-language-models/recognition-simulation-and-refusal-a… · home topics large-language-models article
[ARTICLE · art-136619] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents

A new benchmark called PsyAgentBench, presented in arXiv paper 2609.22090v1, re-ran classic psychology experiments on LLM agents across 41,904 trials and found that apparently human-like effects arise through qualitatively different routes rather than one shared susceptibility. In the Asch conformity paradigm, results ranged from 0 percent under blind framing to 83.3 percent when the paradigm was named on gpt-oss-120B, while anchoring showed exactly zero effect on grounded facts versus near-total effect on invented quantities, and the minimal-group allocation paradigm produced refusal as the primary finding. The authors argue scalar bias-susceptibility scores obscure this structure and instead report replication profiles, formalizing three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection.

by read1 min views2 publishedSep 22, 2026

arXiv:2609.22090v1 Announce Type: new Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.

── more in #large-language-models 4 stories · sorted by recency
── more on @psyagentbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/recognition-simulati…] indexed:0 read:1min 2026-09-22 ·