{"slug": "how-well-synthetic-samples-replicate-public-opinion", "title": "How well synthetic samples replicate public opinion", "summary": "AI-generated synthetic survey respondents differed from real human poll results by an average of 12 percentage points across three American Trends Panel waves, according to a Pew Research Center evaluation. Absolute average error ranged from 11 to 15 percentage points per wave, and no demographic or behavioral subgroup had an average error below 12 points, with Republicans and Republican leaners, Black adults, adults without college attendance, and infrequent or non-internet users performing especially poorly. Pew also reported skewed subgroup results, including that 97% of simulated Hispanic respondents said they were likely to follow the World Cup versus 42% of real Hispanic adults, and 72% of simulated Hispanic adults called teaching Spanish in schools extremely important, more than double the human figure.", "body_md": "Any evaluation of a new polling technology like AI-based surveys starts with a basic question: How well does it replicate the results of high-quality, nationally representative probability polls? Does it accurately depict the shares of Americans who hold various views? And does it work better or worse for particular demographic and behavioral subgroups?\n\nTo evaluate these questions, we used AI respondents to replicate three survey waves previously administered to the [American Trends Panel](https://www.pewresearch.org/the-american-trends-panel/) (ATP). We then calculated the average question-level error rate for each wave, including the error rates for several demographic subgroups.\n\nThese are the three ATP waves we replicated and some of the topics they included:\n\n- **Wave 185** (fielded January 2026): political attitudes and opinions of the Trump administration, military action in Venezuela, annexation of Greenland, conduct of U.S. Immigration and Customs Enforcement officers, and attitudes toward data centers\n- **Wave 190** (fielded March 2026): global politics and international relations, knowledge questions about international affairs and the U.S. Constitution\n- **Wave 192** (fielded April 2026): presidential approval and political attitudes, problems facing the country, military action in Iran, and sleep habits\n\n*This analysis is part of a larger evaluation of AI-generated synthetic samples in public opinion research. Read* [a summary of the main findings](https://www.pewresearch.org/data-labs/2026/09/30/can-ai-stand-in-for-human-survey-takers-not-really/)*and refer to the* [methodology](https://www.pewresearch.org/data-labs/2026/09/30/methodology-silicon-samples/)*for more details on how we conducted our synthetic poll and compared it with human polling results.*\n\nBroadly speaking, the AI survey respondents did not match the views of their human counterparts especially well. Across all three waves, the synthetic polling results differed by an average of 12 percentage points from the estimates produced by the human polls. For individual waves, absolute average error ranged from 11 to 15 percentage points.\n\nThere were also clear errors for all major subgroups in the analysis. For a range of demographic and behavior groups across all three waves, none had an average error of less than 12 points. For several groups, these errors were around 15 points or more.\n\nThe synthetic sample performed especially poorly when it came to replicating the responses of groups including:\n\n- Republicans and Republican leaners\n- Black adults\n- Those who have not attended college\n- Those who use the internet infrequently or not at all\n\n### Synthetic poll performance on specific topics and subgroups\n\nIn addition to the model’s overall poor performance replicating public opinion for different subgroups, we also found that this method produced skewed or biased results for specific questions and subgroups.\n\n#### Views of Hispanic adults toward the World Cup, Spanish classes in schools\n\nAI models are trained on how humans think and act in large part from online content. Because of this, there has long been evidence and concern that they have a [tendency to make assumptions](https://hai.stanford.edu/news/covert-racism-ai-how-language-models-are-reinforcing-outdated-stereotypes) or rely on stereotypes about certain groups of people. We saw several instances of this behavior in our own experiment.\n\nFor example, in a March 2026 ATP survey, fewer than half of real Hispanic adults (42%) said they were at least somewhat likely to follow the World Cup. In contrast, nearly every simulated Hispanic respondent in our AI poll (a full 97%) said they were likely to follow the tournament.\n\nLikewise, the synthetic sample dramatically overestimated Hispanic opinion on the importance of teaching Spanish in schools. According to the synthetic sample, 72% of Hispanic adults believe it is extremely important for schools to offer instruction in Spanish – more than double the rate among real Hispanic adults (32%).\n\n#### Partisan attitudes\n\nIn addition to racial stereotyping, there were also numerous instances of our model misstating the nature and magnitude of partisan differences on issues of the day. Most notably, our AI survey often took views that are held by many Democrats or Republicans – but by no means all of them – and made them appear close to universal.\n\nFor instance, **at least** **80% of Democrats and Democratic leaners in our AI survey** say:\n\n- The fact that some people in the United States have personal fortunes of more than $1 billion is a bad thing for the country (while the actual share is 45%).\n- They would prefer to live in an area where homes are smaller and closer to each other, but with lots of amenities nearby (actually 60%).\n\nSimilarly, **more than 90% of Republicans and Republican leaners in our AI survey** say:\n\n- The police should be allowed to stop and search anyone who fits the general description of a crime suspect (actually 69%).\n- They have very or somewhat favorable views of Israel (actually 58%).\n\nAcross numerous such examples, our AI respondents were far less politically diverse than the actual population.", "url": "https://wpnews.pro/news/how-well-synthetic-samples-replicate-public-opinion", "canonical_source": "https://www.pewresearch.org/data-labs/2026/09/30/how-well-synthetic-samples-replicate-public-opinion/", "published_at": "2026-09-30 17:55:18+00:00", "updated_at": "2026-09-30 18:19:56.581354+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-ethics"], "entities": ["Pew Research Center", "American Trends Panel", "World Cup", "Donald Trump", "U.S. Immigration and Customs Enforcement", "Venezuela", "Greenland", "Iran"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-well-synthetic-samples-replicate-public-opinion", "markdown": "https://wpnews.pro/news/how-well-synthetic-samples-replicate-public-opinion.md", "text": "https://wpnews.pro/news/how-well-synthetic-samples-replicate-public-opinion.txt", "jsonld": "https://wpnews.pro/news/how-well-synthetic-samples-replicate-public-opinion.jsonld"}}