arXiv:2609.05514v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50% of personas fail to express their assigned WVS profiles from the outset, while 2-7% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.
The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
A World Values Survey-grounded simulation of LLM agent societies found that more than 50% of personas fail to express their assigned WVS value profiles from the outset, while only 2-7% drift after repeated conversations, according to the arXiv paper 2609.05514v1. Across roughly 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), the researchers report that simulated value distributions systematically deviate from assigned WVS profiles, and that ablations removing demographic details improve faithfulness for some models without changing the broader trend. The authors conclude that current LLM agents can generate plausible conversations but remain limited proxies for representing and preserving diverse human value profiles over time.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.