cd /news/artificial-intelligence/gemini-4-argon-posts-the-lowest-hall… · home › topics › artificial-intelligence › article
[ARTICLE · art-143066] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Gemini 4 Argon Posts the Lowest Hallucination Rate of Any Top AI Model

Artificial Analysis's AA-Omniscience benchmark found Google's Gemini 4 Argon, launched September 30, posts a 15% hallucination rate — the lowest of any model scoring 45 or above on the firm's Intelligence Index, beating GPT-6 Astra's 51% and GPT-6.1 Sol's 54%. The same benchmark shows Argon's raw accuracy at 50%, five points behind Gemini 3.1 Pro Preview and thirteen points behind GPT-6 Astra's 63%, indicating the model is unusually good at declining to answer rather than being more often correct. On Harvey's Legal Agent Benchmark, Argon scored 19.6% versus GPT-6 Astra's 5.4% and Claude Opus 5.5's 3.8%, and Google gave unrestricted access first to vetted cybersecurity defenders through its Fairwind Program.

by read5 min views1 publishedOct 1, 2026
Gemini 4 Argon Posts the Lowest Hallucination Rate of Any Top AI Model
Image: Startupfortune (auto-discovered)

Is Gemini 4 Argon's hallucination problem actually solved? It scored a 15% hallucination rate on a third-party benchmark, the lowest of any model near the top of the leaderboard, but the same benchmark shows its accuracy at only 50%.

Scroll through r/singularity right now and you'll find threads with hundreds of upvotes insisting Google killed one of AI's oldest problems. "Hallucinations now close to being eliminated," one popular post reads, racking up comments within hours of Gemini 4 Argon's rollout on September 30. That's not quite what happened. What actually happened is narrower, and more interesting than the headline Reddit wrote for it.

Artificial Analysis, the independent benchmarking firm whose Intelligence Index has become a standard reference point for comparing frontier models, ran Argon through its AA-Omniscience benchmark and found a 15% hallucination rate. That's the lowest of any model scoring 45 or above on the firm's Intelligence Index, beating GPT-6 Astra's 51% and GPT-6.1 Sol's 54% by a wide margin. Google's own announcement, posted to its blog on September 30, frames Argon as built for "deep reasoning across complex, long-horizon workflows," with enterprise knowledge work in legal and finance named explicitly as a target use case, alongside software engineering and cybersecurity defense.

Here's the part Reddit skipped. AA-Omniscience doesn't just measure how often a model says something false. It measures what the model does when it doesn't know the answer, specifically the rate of wrong guesses among the responses that weren't simply correct. Argon's 15% figure means it's unusually good at saying "I don't know" instead of making something up. That's a real and useful trait. It is not the same as being right more often. On the same benchmark, Argon's actual accuracy comes in at just 50%, five points behind Gemini 3.1 Pro Preview and thirteen points behind GPT-6 Astra's 63%. Google traded some raw correctness for honesty about its own limits. That's a legitimate engineering choice. It's also not what "hallucinations solved" means to anyone evaluating these systems for real work.

The honesty behind this is where the enterprise case gets interesting. A model that says "I don't know" is far easier to build guardrails around than one that states a wrong answer with total confidence, because a flagged gap can be routed to a human, while a confident fabrication usually isn't caught until it's already in a contract or a chart. That's the practical reason legal and financial workflows are the domains Google is explicitly targeting. On Harvey's Legal Agent Benchmark, Argon scored 19.6%, well ahead of GPT-6 Astra's 5.4% and Claude Opus 5.5's 3.8%, according to data compiled by Artificial Analysis. That's a real gap, and it's the kind of number a law firm evaluating AI tools would actually care about.

Google Restricts Its Most Powerful AI Model to Vetted Cyber Defenders First Google launched Gemini 4 Argon on September 30 and gave unrestricted access first to vetted cybersecurity defenders through its Fairwind Program, not to paying customers. The model leads rivals on most disclosed benchmarks but trails Claude Opus 5.5 on coding-agent tasks, and Google says it's gating the release over the model's offensive cyber... - Google restricts Gemini 4 Argon to cyber defenders - how to access Google's Gemini 4 Argon model

None of this is available to the average person yet, which matters for how fast any of it reaches the industries being promised it. As StartupFortune reported last week, Google rolled Argon out first to a limited group of vetted cyber defenders through its Fairwind Program, alongside the U.S. government's voluntary pre-release access process, rather than to the public. TechCrunch reported the same restricted rollout on September 30, noting Google is withholding a version of Argon with its cybersecurity guardrails removed for internal and vetted-partner use only. Paid API customers and Google AI Ultra subscribers come next, with no public release date committed yet. Nothing's shipped to the public. Introductory API pricing is set at $2 per million input tokens and $10 per million output tokens, with cached tokens priced 95% lower, according to developer platform lead Logan Kilpatrick's announcement on X.

Frankly, "lowest hallucination rate ever measured" is itself a genuinely big deal, and Google earned that headline without needing Reddit's help to inflate it. A frontier model that's two to three times more honest about its own uncertainty than its closest rivals is exactly the kind of incremental, verifiable progress that actually moves enterprise AI forward in legal review, financial analysis, and other high-stakes workflows where a single confident wrong answer can cost more than the model saved. But a 50% accuracy rate is still a coin flip on half the questions AA-Omniscience throws at it. Legal and medical teams evaluating Argon over the next few months won't be asking whether hallucinations are solved. They'll be asking whether a model that admits "I don't know" half the time is more useful than one that occasionally lies to them. For now, that's still an open question, not a solved one.

Also read: Huawei warns of more smartphone price hikes as the memory chip crunch bites • AI coding assistants keep inventing fake packages and hackers are ready • Tencent leases 100,000 AI chips from Oracle in a $7 billion five year deal

This article is posted in AI News, check it out for more related stories.

Join the discussion #

Open in the community → Almost there. Sign in and your reply posts straight away.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-4-argon-posts…] indexed:0 read:5min 2026-10-01 · —