{"slug": "beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins", "title": "Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins", "summary": "A new arXiv study (2608.20344v1) finds that the main constraint in LLM-based digital twins is not information volume but how persona information is structured, with a hand-crafted schema (BDE: Background, Decision procedure, Evaluation) improving predictive accuracy by +1.91 percentage points over raw transcripts on a homogeneous benchmark (Twin-2K-500), but failing on heterogeneous tasks. An automatic structure-discovery pipeline, where an LLM iteratively refines task-specific structures, restores performance on a benchmark of 13 diverse sub-studies, improving mean accuracy by +1.91 percentage points over the raw transcript baseline.", "body_md": "arXiv:2608.20344v1 Announce Type: new\nAbstract: LLM-based \"digital twins\" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck.\nIn this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline.\nTo address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema.\nOverall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.", "url": "https://wpnews.pro/news/beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins", "canonical_source": "https://arxiv.org/abs/2608.20344", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 04:14:03.452416+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["arXiv", "Twin-2K-500", "gpt-5.4-mini", "Qwen3-8B"], "alternates": {"html": "https://wpnews.pro/news/beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins", "markdown": "https://wpnews.pro/news/beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins.md", "text": "https://wpnews.pro/news/beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins.txt", "jsonld": "https://wpnews.pro/news/beyond-raw-transcripts-structured-persona-extraction-for-llm-based-digital-twins.jsonld"}}