cd /news/artificial-intelligence/beyond-raw-transcripts-structured-pe… · home topics artificial-intelligence article
[ARTICLE · art-108243] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

A new arXiv study (2608.20344v1) finds that the main constraint in LLM-based digital twins is not information volume but how persona information is structured, with a hand-crafted schema (BDE: Background, Decision procedure, Evaluation) improving predictive accuracy by +1.91 percentage points over raw transcripts on a homogeneous benchmark (Twin-2K-500), but failing on heterogeneous tasks. An automatic structure-discovery pipeline, where an LLM iteratively refines task-specific structures, restores performance on a benchmark of 13 diverse sub-studies, improving mean accuracy by +1.91 percentage points over the raw transcript baseline.

read1 min views1 publishedAug 24, 2026

arXiv:2608.20344v1 Announce Type: new Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-raw-transcrip…] indexed:0 read:1min 2026-08-24 ·