arXiv:2609.13995v1 Announce Type: new Abstract: Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.
Synthetic Data in Marketing Research: How to Evaluate and When to Trust
A new arXiv paper (2609.13995v1) argues that the debate over synthetic data in marketing research should shift from whether LLM-generated respondents work to when they work, distinguishing three types of synthetic data: ungrounded LLM responses, segment-level personas, and individual-level digital twins. The authors propose an ex-ante answerability diagnostic — the R² of a random forest predicting twin outputs from the data used to construct the twins — and report that across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R² above 0.7 raises the mean twin-human individual-level correlation by 15% and cuts the share of poorly answered questions from 25.9% to 4.3%. The paper also introduces the "forgotten question problem," in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data, and finds embedding similarity and experienced-researcher judgment provide correlated but weaker screens.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.