cd /news/large-language-models/simulating-social-attitudes-with-llm… · home topics large-language-models article
[ARTICLE · art-57982] src=aclanthology.org ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Simulating Social Attitudes with LLMs: Accuracy, Demographic Effects, and Refusal Behavior in the Sensitive Domain of Suicide Prevention

A study by Cristina J. Perez, Michael P. Vasquez Jr, Philippe Giabbanelli, and Patrick Y. Wu found that large language models (LLMs) have a mean absolute error of 23 percentage points when simulating public attitudes toward suicide prevention policies, based on 811,560 prompts using 32 questions from seven U.S. surveys (2023-2025). The researchers tested GPT-5 Nano, DeepSeek V3.2, Meta Llama 3.1 8B, and Mistral Small 24B, finding that model choice affects accuracy more than prompt framing, while refusal behavior varies sharply across models and prompt designs.

read2 min views17 publishedJul 13, 2026
Simulating Social Attitudes with LLMs: Accuracy, Demographic Effects, and Refusal Behavior in the Sensitive Domain of Suicide Prevention
Image: Aclanthology (auto-discovered)

Simulating Social Attitudes with LLMs: Accuracy, Demographic Effects, and Refusal Behavior in the Sensitive Domain of Suicide Prevention

[Cristina J. Perez](/people/cristina-j-perez/unverified/),
[Michael P. Vasquez Jr](/people/michael-p-vasquez-jr/unverified/),
[Philippe Giabbanelli](/people/philippe-giabbanelli/),
[Patrick Y. Wu](/people/patrick-y-wu/unverified/)
Abstract

Large language models (LLMs) are increasingly used to simulate public opinion, yet their validity in sensitive policy domains remains underexplored. We evaluate whether LLMs can reproduce attitudes toward suicide prevention policies using 32 questions drawn from seven nationally representative U.S. surveys (2023-2025). We systematically vary demographic conditioning (race/ethnicity, gender, age, education, income, party), prompt framing (direct elicitation, respondent embodiment, specialist embodiment), and model architecture (GPT-5 Nano, DeepSeek V3.2, Meta Llama 3.1 8B, Mistral Small 24B). Across 811,560 prompts, the mean absolute error—the average gap between predicted and human response distributions—is 23 percentage points. We also find that LLM responses to demographic-conditioned prompts diverge substantially from prompts without demographic information. In short, what distribution LLMs draw on when generating responses to sensitive polling questions remains unclear. Model choice matters more than framing for accuracy, whereas refusal behavior varies sharply across models and prompt designs. Our findings highlight the limitations of LLMs for social simulation in the context of sensitive topics.- Anthology ID:

- 2026.nlpcss-1.12
- Volume:
[Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science](/volumes/2026.nlpcss-1/)- Month:
  • July
  • Year:
  • 2026
  • Address:
  • San Diego
- Editors:
[Dallas Card](/people/dallas-card/),[Anjalie Field](/people/anjalie-field/),[Katherine Keith](/people/katherine-keith/),[Julia Mendelsohn](/people/julia-mendelsohn/)- Venues:
[NLP+CSS](/venues/nlpcss/)|[WS](/venues/ws/)- SIG:
- Publisher:
  • Association for Computational Linguistics
- Note:
- Pages:
  • 176–189
- Language:
- URL:
[https://aclanthology.org/2026.nlpcss-1.12/](https://aclanthology.org/2026.nlpcss-1.12/)- DOI:
[10.18653/v1/2026.nlpcss-1.12](https://doi.org/10.18653/v1/2026.nlpcss-1.12)- Cite (ACL):
── more in #large-language-models 4 stories · sorted by recency
idiallo.com · · #large-language-models
BI Slop
── more on @gpt-5 nano 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/simulating-social-at…] indexed:0 read:2min 2026-07-13 ·