{"slug": "recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm", "title": "Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents", "summary": "A new benchmark called PsyAgentBench, presented in arXiv paper 2609.22090v1, re-ran classic psychology experiments on LLM agents across 41,904 trials and found that apparently human-like effects arise through qualitatively different routes rather than one shared susceptibility. In the Asch conformity paradigm, results ranged from 0 percent under blind framing to 83.3 percent when the paradigm was named on gpt-oss-120B, while anchoring showed exactly zero effect on grounded facts versus near-total effect on invented quantities, and the minimal-group allocation paradigm produced refusal as the primary finding. The authors argue scalar bias-susceptibility scores obscure this structure and instead report replication profiles, formalizing three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection.", "body_md": "arXiv:2609.22090v1 Announce Type: new \nAbstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.", "url": "https://wpnews.pro/news/recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm", "canonical_source": "https://arxiv.org/abs/2609.22090", "published_at": "2026-09-22 04:00:00+00:00", "updated_at": "2026-09-22 04:25:05.275980+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-agents", "ai-safety"], "entities": ["PsyAgentBench", "gpt-oss-120B", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm", "markdown": "https://wpnews.pro/news/recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm.md", "text": "https://wpnews.pro/news/recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm.txt", "jsonld": "https://wpnews.pro/news/recognition-simulation-and-refusal-a-contamination-aware-study-of-classic-in-llm.jsonld"}}