cd /news/artificial-intelligence/just-facts-study-tests-political-cha… · home topics artificial-intelligence article
[ARTICLE · art-95654] src=letsdatascience.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Just Facts Study Tests Political Chatbot Factuality

Just Facts, a nonprofit research group, tested paid versions of ChatGPT, Gemini, Grok, and Claude with 100 political questions and found that ChatGPT, Gemini, and Claude scored lower on prompts tied to left-coded falsehoods, while Grok showed the reverse pattern, according to a Fox News report. The study also found that roughly half of the cited AI sources were illegitimate, meaning they did not support the associated claim or did not exist. Jim Agresti, president and co-founder of Just Facts, told Fox News the exercise was intended to examine misinformation rather than political bias alone.

read3 min views1 publishedAug 13, 2026
Just Facts Study Tests Political Chatbot Factuality
Image: Letsdatascience (auto-discovered)

Just Facts tested paid versions of ChatGPT, Gemini, Grok and Claude with 100 political questions designed to elicit falsehoods from both ideological directions in a study reported by Fox News. The study found ChatGPT, Gemini and Claude scored lower on prompts tied to left-coded falsehoods, while Grok showed the reverse pattern and roughly half of cited AI sources were illegitimate.

Just Facts tested paid versions of ChatGPT, Gemini, Grok, and Claude with 100 questions on politically contested topics, according to Fox News. The questions covered immigration, abortion, climate change, elections, crime, gun control, and COVID-19, and were designed to prompt false statements associated with both the political left and right.

Fox News reported that ChatGPT correctly answered 94% of questions designed to elicit right-coded falsehoods and 75% of those designed to elicit left-coded falsehoods. Gemini scored 91% and 76%, respectively, while Claude scored 91% and 81%. Grok showed the opposite pattern, with 73% on right-coded prompts and 84% on left-coded prompts, according to the outlet's account of the study.

The report also said that roughly half of the sources supplied by the chatbots were illegitimate, meaning they either did not support the associated claim or did not exist. Jim Agresti, president and co-founder of Just Facts, told Fox News that the exercise was intended to examine misinformation rather than political bias alone.

What the evaluation measures

The reported results concern responses to a deliberately adversarial set of political prompts, not a general benchmark of each system's overall factual accuracy. That distinction matters: accuracy rates can vary substantially with prompt wording, question selection, model version, system instructions, retrieval settings, and how a study defines a falsehood and a valid citation.

For ML teams deploying general-purpose assistants in high-stakes information settings, the source-quality finding is as operationally important as the answer-level scores. A fluent answer with a citation can still fail verification if the cited page is irrelevant, unavailable, or fabricated. Comparable evaluations across the sector commonly require separate measurement of claim accuracy, citation entailment, source existence, and susceptibility to user pressure. Those dimensions help distinguish a model that states an incorrect answer from one that compounds the error by presenting unsupported evidence.

Limits for practitioners

Fox News reported the study's topline scores and methodology, but the available report does not establish how results would transfer to other prompt sets or deployment configurations. Teams assessing political or public-policy use cases should therefore reproduce tests against their own model versions, retrieval pipelines, and approval workflows rather than treating a single adversarial audit as a complete safety assessment.

Key Points #

  • 1Fox News reported asymmetric accuracy scores across ideological prompt sets, making evaluation design central to interpreting the chatbot comparisons.
  • 2The study's reported source failures matter because unsupported citations can make incorrect generated answers appear more credible to users.
  • 3Comparable high-stakes deployments generally benefit from separately testing answer accuracy, citation validity, and resistance to adversarial conversational pressure.

Scoring Rationale #

The report covers factuality and citation reliability in four widely used AI assistants, issues directly relevant to teams deploying LLMs for public-facing information. Its practical significance is tempered by the politically focused, adversarial question set and the need for independent replication across model configurations.

Sources #

Primary source and supporting public references used for this report.

Practice interview problems based on real data

1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.

Try 250 free problems

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @just facts 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/just-facts-study-tes…] indexed:0 read:3min 2026-08-13 ·