cd /news/artificial-intelligence/rethinking-the-evaluation-and-optimi… · home topics artificial-intelligence article
[ARTICLE · art-105425] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

Researchers from arXiv introduced the subjectivity coefficient, an entropy-based measure, to address the unreliability of accuracy-based evaluation and hard-label training for LLM-based social simulation, proposing Subjectivity-Adaptive soft-Label Training (SALT) which pools outputs from semantically nearby inputs into soft labels. They also constructed SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions, demonstrating SALT's advantages in realistic settings where only single observations are available.

read1 min views3 publishedAug 21, 2026

arXiv:2608.19689v1 Announce Type: new Abstract: LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and trains the LLM to reproduce this hard label. However, human behavior is inherently subjective: the same person in the same situation may reasonably act differently, so an observed response is only one draw from an underlying response distribution, rendering accuracy-based evaluation unreliable and hard-label training misleading. To address these problems, we first introduce the subjectivity coefficient, an entropy-based quantity distinguishing objective tasks such as coding from subjective ones such as social simulation, and use it to systematically analyze how accuracy-based evaluation and hard-label training fail as subjectivity grows. Based on the subjectivity coefficient, we propose Subjectivity-Adaptive soft-Label Training (SALT): it pools observed outputs from semantically nearby inputs into soft distributional labels, with an aggregation radius adapted to the estimated subjectivity of each input; in the near-objective limit the neighborhood shrinks, so SALT naturally falls back to standard single-label training. Moreover, since existing datasets record only single observed responses and cannot support distributional evaluation, we construct SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions. Since real-world data typically provide only a single observation per input, our experiments train models from single observed outputs while evaluating them against the full response distributions, verifying feasibility in realistic settings. Results on SUBJSIM demonstrate the advantages of our method.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rethinking-the-evalu…] indexed:0 read:1min 2026-08-21 ·