As artificial intelligence becomes more advanced, there is growing interest in using it to stand in for real respondents in public opinion surveys. Put simply: Instead of contacting large numbers of people and asking them what they think, pollsters can ask an AI model to predict how those people would have answered a certain question.
Pew Research Center has long sought to better understand new developments in public opinion research, and we wanted to learn more about how this AI-based polling works.1So we ran an experiment. We developed a state-of-the-art process for fielding AI surveys and compared their results with three recent survey waves from our American Trends Panel (ATP) taken by human respondents at roughly the same time.
This experiment taught us that, at this time, AI models are not an adequate replacement for traditional polling on topics of broad public importance. Our AI-generated survey results struggled to accurately reproduce the findings from high-quality public opinion polls in a variety of ways.
Here are some of the main issues we encountered:
Results of AI polls differed – often by quite a lot – from results of human surveys
Across nearly 300 individual survey questions, the estimates produced using our AI respondents differed from their human counterparts by an average of 12 percentage points.
The average difference exceeded 15 points on around 28% of questions we asked. And it was not uncommon to see differences exceeding 20, 30 or even 40 percentage points on individual answers to some questions.
Read more about how well synthetic samples replicate public opinion.
Major misses on ‘timely and topical’ political issues
Our AI survey missed the mark on numerous questions about current events in early 2026. Among other things:
- It overstated the share of Americans who approved of President Donald Trump’s job performance at the time, despite the availability of long-standing trend data.
- It greatly underestimated the share of Republicans who think it’s acceptable for immigration officers to wear face coverings, resulting in a large overall error.
- It missed badly on a variety of questions about data centers, from general awareness to views about their impact.
Often, these misses were unpredictable. On questions about cost-of-living concerns, our model somewhat overestimated the shares of Americans who are concerned about the cost of healthcare and consumer goods while greatly underestimating the share who are worried about the price of electricity.
Read more about synthetic surveys and “timely and topical” questions.
AI consistently avoids some answers but piles onto others
Human opinion is extremely diverse, but our AI respondents often avoided certain answers altogether. Nearly half the questions we asked had at least one answer choice that was not selected by a single AI-generated respondent.
These errors can distort true public opinion. For instance, the model (accurately) estimated that a majority of U.S. adults support legal abortion. But while a number of Americans think abortion should be legal in all cases (23%) or illegal in all cases (11%), the AI poll underestimated these opinions.
AI polling results can lean on racial or partisan stereotypes
In many cases, our AI poll took beliefs or attitudes that are reasonably common among a particular subgroup and portrayed them as nearly ubiquitous. For example, our AI poll would indicate that:
- 97% of Hispanic adults are at least somewhat likely to follow the World Cup (while the actual share is 43%).
- 95% of Republicans and Republican-leaning independents have a very or somewhat favorable view of Israel (actually 58%).
- 86% of Democrats and Democratic leaners think billionaires are a bad thing for the country (actually 45%).
Read more about how well synthetic polls replicate the diversity and distribution of public opinion.
The model thinks we know more than we actually do
Do you know which right is protected by the First Amendment to the U.S. Constitution? Or what the focus of the NATO alliance is? Our surveys have found that fewer than six-in-ten Americans can correctly answer these questions. But our AI poll estimated that nearly every member of the public knows the answers.
More broadly, our AI respondents didn’t like to admit to uncertainty. Across all the questions on our three surveys where a “not sure” option was offered, human panelists were around four times as likely as the AI model to choose it.
Read more about how synthetic respondents express certainty and factual knowledge.
The model you use can change the answers you get
We found that different AI models can paint a different picture of the public mood – even when everything else about the survey is the same. In a test comparing OpenAI’s GPT-5.1 and Anthropic’s Claude Opus 4.6 on a subset of questions, the GPT estimates described an American public that has more extreme opinions on a variety of topics than they actually do, while the Opus estimates described a public that is more middle-of-the-road than in reality. A reader of either result would be misled, but in different directions.
Read more about how synthetic polling results change based on the AI model.
The bottom line: Survey-taking is best left to humans
Our biggest takeaway from this exercise is that AI polling is not a replacement for rigorously surveying real humans. It’s not just that the AI results differ from human results, although that is certainly true. The bigger story is that these results differ in ways that are often unpredictable, and they are highly subject to factors like unforeseen real-world events or the choice of model used.
To be sure, this is a fast-evolving field with a great deal of academic and applied research effort behind it. There are also many other ways to use AI to improve the polling process that stop short of replacing human respondents. For instance, AI can be used to categorize real responses to open-ended questions or write the code used to analyze the survey results – the Center is using AI tools in these contexts and will continue to do so. But at the end of the day, we see no substitute for rigorous, multimode, probability-based surveys that allow real members of the public to speak their minds on issues of importance.