Four AI models publicly calling out a fifth for citing human hallucination statistics instead of AI hallucination statistics is not a bug story. It is the pitch.
Suprmind, founded by Radomir Basta, is built around a simple premise: comparing AI outputs works better when the models are all in the same room, reading each other's answers, rather than scattered across five separate browser tabs. The platform lets a user run frontier models side by side in one shared thread instead of copying a question into ChatGPT, then Claude, then Grok, then Perplexity, then Gemini, and trying to reconcile five different answers by hand.
Why one thread beats five tabs #
The obvious benefit of a shared thread is speed. Nobody has to retype a question five times or keep five tabs straight. But according to Basta, the more important benefit is not about saving time at all. It is about what happens when the models can see each other's work.
Basta described it directly: running multiple AIs in the same thread does more than save time, because AI models are known to hallucinate, misstate, and fabricate data. When five models are all reading everybody else's answers, and arguing and disagreeing about the reasoning behind them, that back and forth filters out a meaningful amount of false data and flawed reasoning as part of the normal conversation flow. The correction happens inside the same window where the question was asked, rather than requiring the user to separately notice something looked off and go check a second source.
He offered a concrete example rather than a general claim. Basta asked the group for hallucination statistics. Perplexity returned data that read as strong on the surface, data he says he would have taken at face value if it had come back alone. Grok, answering next in the same thread, flagged the problem: the statistics Perplexity had surfaced were not about AI hallucinations at all, but about hallucinations caused by mental illness in humans. The other models in the thread followed Grok's lead and called out the mismatch as well, catching an error a single-platform workflow would likely have missed entirely.
Ilya Sutskever's SSI Has Raised $8 Billion and Shipped Nothing At All Ilya Sutskever's Safe Superintelligence has raised about $8 billion, including a fresh $5 billion from Nvidia, without shipping a single product. Sutskever says the industry's scaling era is over and SSI is betting on something else entirely, but two years in, the company still hasn't shown the world what that something is. - how to price a SaaS product for enterprise - cold email template that gets replies from investors
What makes that example useful as a pitch is not just that an error occurred, but where it was caught. Basta's own framing of it, that he would have accepted Perplexity's answer at face value in isolation, is the point: a fluent, confidently formatted statistic is not, on its own, evidence that it is correct. The mismatch was only visible because a second model was reading the same claim in the same thread and had reason to compare it against its own understanding of the question.
The case for cross-checking as a feature #
That kind of exchange is the core argument for running models together instead of apart. A single AI answer, however confident it sounds, is still one model's output with no built-in check on it. Suprmind's structure turns the comparison itself into the check: when a model gets something wrong, the others in the thread are positioned to notice, because they are seeing both the question and the competing answers in real time.
Basta backs this pattern with data beyond individual anecdotes. Suprmind has published its own research on how much frontier models diverge from each other when answering the same prompts, drawn from production data generated by active users of the platform rather than a synthetic benchmark. That work is published at suprmind.ai/hub/multi-model-ai-divergence-index, and it exists precisely because divergence between models is not rare. It is common enough that a platform built around users interacting with several models at once had reason to study it directly. Basing that research on real usage rather than a constructed test set matters for the same reason the hallucination-statistics anecdote does: the divergence Suprmind is measuring is the same divergence that shows up in ordinary, everyday questions asked by people who are not trying to stress-test the models at all.
For anyone relying on a single AI platform for research, writing, or decision support, that divergence is the underlying risk Suprmind is positioned against. A wrong answer that goes unchallenged looks identical to a right one until someone happens to check it elsewhere. Suprmind's answer is to build the checking into the workflow itself rather than leaving it to the user to catch after the fact, or not catch at all. That shifts the burden of verification away from the individual user's vigilance, which is inconsistent and easy to skip under time pressure, and toward a structural feature of the platform that runs the same way every time a question is asked, regardless of whether the user happens to be paying close enough attention to notice a subtle error.
Built for people already comparing AI outputs #
Suprmind is aimed squarely at people who already do this kind of comparison manually: developers, researchers, and anyone using AI tools heavily enough to have noticed that different models give meaningfully different answers to the same question. Rather than asking those users to keep juggling separate chats and reconciling the differences themselves, the platform puts the comparison, and the disagreement, in one place where it can be watched as it happens. For that audience, the pitch is less about discovering that models disagree, which they likely already suspect from experience, and more about no longer having to be the one who manually cross-references five answers every time a question matters enough to double-check.
That includes moments Basta describes as genuinely entertaining to watch, such as multiple models pushing back on one another over outdated or mismatched data in real time. But the entertainment value is really a byproduct of the mechanism doing its job: models checking each other is what surfaces the error in the first place.
Suprmind's multi-AI platform puts that shared-thread approach at the center of the product rather than treating it as a side feature.
AIUC Wants To Insure Your AI Agents Before They Go Rogue AIUC, a startup founded by Anthropic's first go-to-market hire Rune Kvist and former METR COO Rajiv Dattani, runs AI agents through roughly 5,000 adversarial tests and issues 100-page audit reports before enterprises deploy them. Cursor, Lovable, Harvey and ElevenLabs are already customers, and the business is landing amid a wave of real... - how to insure AI agents before deployment - AI agent testing and security audit requirements
Join the discussion #
Open in the community → Almost there. Sign in and your reply posts straight away.