arXiv:2610.10988v1 Announce Type: new Abstract: As agentic systems gain commercial popularity, user simulators increasingly serve as measurement instrument for their evaluation. However, the fidelity of simulated users in comparison to real human users is generally low, and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. We author personas as concrete stylistic rates, which lets us transfer two established, model-free instruments -- authorship-verification stylometry and lexicon-based content analysis -- as fidelity diagnostics. We A/B-test the sociolinguistic schema against a flat descriptive baseline across five task-oriented customer-service agents. Results show that the sociolinguistic schema improves both stylistic adherence and stylometric distinguishability for most of the tested models, with a caveat that persona style fidelity does not necessarily equal persona "naturalness". We argue that a sociolinguistic approach to persona design is a promising path towards more diverse and representative user personas, and that these metrics are most valuable in an error-attribution analysis, localizing where fidelity breaks down. This is a first step towards interventions that move user simulations closer to faithful renderings of diverse and variable linguistic outputs.
Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation
A pilot study on arXiv (2610.10988v1) reports that authoring user-simulation personas as concrete stylistic rates and measuring fidelity with model-free stylometry and lexicon-based content analysis improved stylistic adherence and stylometric distinguishability for most tested models across five task-oriented customer-service agents. The sociolinguistic schema was A/B-tested against a flat descriptive baseline, with the authors noting that persona style fidelity does not necessarily equal persona naturalness. The authors argue the approach is a first step toward more diverse and representative user personas and that the metrics are most valuable in error-attribution analysis localizing where fidelity breaks down.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.