{"slug": "how-synthetic-polling-results-change-based-on-the-ai-model", "title": "How synthetic polling results change based on the AI model", "summary": "Pew Research Center found that synthetic survey samples generated by OpenAI's GPT-5.1 and Anthropic's Claude Opus 4.6 both failed to replicate human polling results from its American Trends Panel, with Opus averaging an 11.4 percentage-point absolute error versus 13.3 points for GPT. Simulating 6,700 real respondents on a questionnaire fielded in late January 2026, the two models diverged sharply from each other and from actual public opinion — GPT estimated 100% of Americans were dissatisfied with the country's direction, while Opus missed by roughly 30 percentage points on whether clear solutions exist to major national problems. Pew concluded that neither model truly reflected actual public opinion, and no topic category emerged as a consistent strength for either model.", "body_md": "A key feature of public opinion polling is its consistency and predictability. Although the individual participants may be different, two identical surveys fielded at the same time using the exact same methods should produce largely similar results within known boundaries.\n\nBut there is no guarantee of this when the “participants” taking the surveys are AI models. To explore how the choice of model can affect the results of the synthetic survey, we used both OpenAI’s GPT-5.1 and Anthropic’s Claude Opus 4.6 to simulate 6,700 real respondents from our [American Trends Panel](https://www.pewresearch.org/the-american-trends-panel/) (ATP) and generate synthetic survey data using a questionnaire we originally fielded in late January 2026. We then compared the models’ results against each other, as well as against the results from the human poll.\n\nLooking at the survey average absolute error, neither of the models we tested replicated the human panel results especially well. Opus performed slightly better, with an average absolute error of 11.4 percentage points (compared with 13.3 percentage points for GPT).\n\nTo be sure, newer models and future innovations in synthetic sample construction could result in smaller overall error figures. The real story here is that **each synthetic sample painted a very different picture of the American public on multiple dimensions, and neither truly reflected actual public opinion.**\n\n*This analysis is part of a larger evaluation of AI-generated synthetic samples in public opinion research. Read* [a summary of the main findings](https://www.pewresearch.org/data-labs/2026/09/30/can-ai-stand-in-for-human-survey-takers-not-really/)*and refer to the* [methodology](https://www.pewresearch.org/data-labs/2026/09/30/methodology-silicon-samples/)*for more details on how we conducted our synthetic poll and compared it with real survey results.*\n\nThe questions and topics on which GPT most closely mirrored actual public sentiment were almost entirely different from those on which Opus was the better performer. Still, there were [no topics or particular categories of question](#_Appendix_D:_Additional) that stood out as a strength of one model over the other.\n\nHere are examples of ways in which the two models provided results that were different from each other – and also from our human poll.\n\n### Example 1: Public views of politics and democracy\n\nOur January ATP survey included a set of questions about broad political attitudes, like overall satisfaction in the way things are going in the United States, which ideological “side” is losing more often, the impact of voting, and whether there are clear solutions to the problems facing the country. These are all fairly long-standing trend questions with ample history to draw on, yet the two models produced wildly inconsistent results.\n\nFor instance, the Opus poll was almost exactly in line with actual public sentiment when it came to the share of Americans who are dissatisfied with the way things are going in the country today; the share who feel like their “side” has been losing more than winning in politics; and the share who say voting gives people like them some say in how the government runs things. Each of these opinions is held by a majority of Americans, but they are far from ubiquitous.\n\nBy contrast, our GPT poll depicts these views as quite literally universal – for instance, estimating that 100% of Americans are dissatisfied with the way things are going in the country.\n\nOn the share of Americans who agree that there are clear solutions to most big issues facing the country today, GPT more closely mirrors the views of the public, while Opus differs from true public opinion by roughly 30 percentage points.\n\n### Example 2: Abortion attitudes\n\nAbortion is an example of an issue on which neither model produced results that accurately reflect public sentiment.\n\nIn our human poll, around six-in-ten Americans said abortion should be legal in all or most cases, while around four-in-ten said it should not be legal. Opus produced the same general split, but greatly underestimated the share of Americans who hold very strong views on abortion:\n\n- 23% of U.S. adults think abortion should be *legal in all cases* ; our Opus poll estimated that share at just 13%.\n- 11% think abortion should be *illegal in all cases* ; Opus estimated that 0% of the public holds this view.\n\nBy contrast, GPT produced overall estimates that imply the country is about evenly split on the legality of abortion, with the view that it should be illegal slightly more common.\n\nAll told, neither synthetic poll illustrates where the country stands on this issue – and they fail to capture the full scope of public sentiment in different ways. Here, the choice of model alone has the potential to reshape the narrative around public opinion of abortion.\n\n### Example 3: Partisan attitudes\n\nThe two models also painted very different pictures of the views held by Republicans and Republican leaners. Synthetic estimates for Republicans and those who voted for President Donald Trump in 2024 make for particularly striking examples of how using a different model can lead to very different conclusions about public opinion within a certain group.\n\nOur January 2026 ATP survey found that Trump’s overall approval rating among his 2024 voters was at 84%, with 65% approving very strongly and 19% approving not so strongly. Both of our synthetic polls overestimated his approval rating – Opus by 15 points and GPT by 13 points. But beyond the topline figures, each synthetic poll presented a different conclusion about this voter base:\n\n- Trump voters in the GPT poll consisted almost entirely of those who approve very strongly.\n- Trump voters in the Opus poll were evenly split between very strong and not so strong approval.\n\nAcross other questions in the survey, GPT estimates tended to support a view of Republicans as fervent Trump supporters with highly polarized partisan views. By contrast, the Opus poll would indicate that far more Republicans hold mixed or moderate views. Some questions where the two models present opposing views of Republican attitudes include:\n\n- Whether Republicans in Congress have an obligation to support Trump’s policies and programs because he is a Republican president\n- Whether or not Trump should work with Democratic congressional leaders\n- Whether China is an enemy, competitor or partner to the United States.\n- Whether or not the U.S. should send ground troops to Venezuela\n\nThey also present very different estimates of public sentiment toward prominent political figures. In our human poll, 74% of Republicans and Republican-leaning independents expressed a favorable view of Health and Human Services Secretary Robert F. Kennedy Jr. Opus estimated his favorability among Republicans at 85%, while the GPT poll estimated their opinion to be 63% *unfavorable*.\n\n### Example 4: Magnitude of opinion and “extreme” answer options\n\nThe example of abortion notwithstanding, the GPT and Opus polls often agreed that the public generally leans in one direction or another on various issues. But often, the two models disagreed on the intensity of those attitudes.\n\nFor instance, 68% of U.S. adults in our ATP survey said they are concerned about the price of gasoline, split evenly between those who are very or only somewhat concerned. GPT and Opus both (incorrectly) estimated that nearly every American is concerned about this issue. But the GPT sample lumps most of the public into the “very concerned” category, while the Opus sample leans heavily toward “somewhat concerned.”\n\nAs it turns out, this kind of pattern appeared consistently throughout the survey. GPT tended to favor response options at the far ends of scales, such as “extremely” or “not at all.” By contrast, Opus tended to favor “middle” options like “somewhat,” “about right” or “neither.” This was broadly true regardless of what the scale was or what the question was about.\n\nFor all questions in the survey with response options arranged in a clear order:\n\n- Human respondents chose a “middle” option 45% of the time, on average.\n- GPT respondents did so 31% of the time.\n- Opus respondents did so 56% of the time.\n\nPut simply: GPT tended to depict Americans as having more extreme opinions than they actually do, while Opus tended to depict them as being more middle-of-the-road than they are in reality. It’s important to note that changes to how synthetic survey data is generated and future updates to the GPT and Opus models themselves could lead to different patterns. However, it is clear that the choice of model alone is enough to yield large, systematic differences between synthetic samples that were otherwise built in exactly the same way.[3](#fn-574146-3)\n\nWhen compared with the data from our real panelists, neither synthetic poll was able to provide a reasonably similar snapshot of the views of the American public. However, given that Claude Opus 4.6 had better performance *on average*, this is the model we selected to generate additional synthetic samples.", "url": "https://wpnews.pro/news/how-synthetic-polling-results-change-based-on-the-ai-model", "canonical_source": "https://www.pewresearch.org/data-labs/2026/09/30/how-synthetic-polling-results-change-based-on-the-ai-model/", "published_at": "2026-09-30 17:55:21+00:00", "updated_at": "2026-09-30 18:19:43.369317+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "generative-ai"], "entities": ["Pew Research Center", "OpenAI", "GPT-5.1", "Anthropic", "Claude Opus 4.6", "American Trends Panel"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-synthetic-polling-results-change-based-on-the-ai-model", "markdown": "https://wpnews.pro/news/how-synthetic-polling-results-change-based-on-the-ai-model.md", "text": "https://wpnews.pro/news/how-synthetic-polling-results-change-based-on-the-ai-model.txt", "jsonld": "https://wpnews.pro/news/how-synthetic-polling-results-change-based-on-the-ai-model.jsonld"}}