# Why AI Benchmarks Are Total BS (And How OpenAI and Anthropic Use Them to Trick You)

> Source: <https://au.pcmag.com/ai/119880/why-ai-benchmarks-are-total-bs-and-how-openai-and-anthropic-use-them-to-trick-you>
> Published: 2026-09-12 12:00:00+00:00

When new [AI](/ai) models hit the scene, companies often cite their scores on benchmarks such as BioMysteryBench, GeneBench, and Terminal-Bench as evidence of progress. But what do these tests evaluate, and do the scores even matter? I dove into the wild world of AI benchmarks through the lens of [GPT-5.6](/ai/118792/return-of-the-king-gpt-56-pulls-ahead-of-fable-5-and-grok-45-in-my-latest-testing), [Fable 5.1](/ai/119733/fable-51-fixed-its-worst-flaws-but-im-still-sticking-with-opus), and [Opus 5](/ai/119350/i-vibe-coded-an-app-with-opus-5-to-get-better-at-the-finals-and-its-working) and found out they’re inconsequential at best and misleading at worst. Moreover, I don’t expect them to improve anytime soon. Buckle up for a deep dive into the numbers and problems.

## Unlike Hardware, You Can't Benchmark Intelligence

No individual metric or test can tell you everything about a tech product. For example, a GPU's [3DMark score](/graphic-cards/63911/how-we-test-graphics-cards) alone isn't enough to tell you how well it will perform in your use case. Maybe your favorite game relies much more heavily on your [CPU](/processors/65778/the-best-cpus) than on your GPU, which means the differences between graphics cards aren't as relevant. However, these benchmarks aren’t useless. If you look at [3DMark’s leaderboards](https://www.3dmark.com/hall-of-fame-2/), there’s a reason why the [RTX 5090](/graphics-cards/109480/nvidia-geforce-rtx-5090-founders-edition) is at the top: It’s indisputably the most powerful consumer-level GPU available today.

AI benchmarks don’t work the same way. For example, [according to Anthropic](https://www.anthropic.com/news/claude-opus-5), Opus 5 scores an impressive 43.3% in the coding-focused Frontier-Bench, while GPT-5.6 earns a paltry 34.4%. Nonetheless, I prefer GPT-5.6 over Opus 5 for my vibe coding projects, thanks to its more consistent performance and better outputs. 

It's unthinkable that an [RTX 5080](/graphics-cards/109498/nvidia-geforce-rtx-5080-founders-edition) would outperform an RTX 5090 after looking at their benchmark scores, but it's perfectly reasonable to prefer one AI model over another for a given task even if its benchmark score is lower. The latter shouldn't be possible if the benchmark is actually valuable.

Moreover, improvements on AI benchmarks rarely translate to appreciable advantaged in real-world usage, whereas you are far more likely to notice an uptick in gaming frame rates or a reduction in project rendering times with a more powerful GPU.

## The Numbers Are Ghostly, Inconsistent, and Unverifiable

When [OpenAI announced GPT-5.6](https://openai.com/index/previewing-gpt-5-6-sol/), it included a graph of the model’s performance on Terminal-Bench to demonstrate its coding skills, saying, “GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1.” Terminal-Bench 2.1, according to its [GitHub page](https://github.com/harbor-framework/terminal-bench-2-1), “is a collection of benchmarks for measuring agents' abilities to complete valuable and complex tasks in container environments.” As you might expect, the graph (shown below) has GPT-5.6 Sol, GPT-5.6’s most powerful variant, at the top with a score of 91.9%. So, you might reasonably expect that GPT-5.6 Sol is the best AI model for doing the kinds of tasks Terminal-Bench 2.1 tests. However, that expectation just doesn’t align with the far more complex reality.

If you visit [the Terminal-Bench 2.1 leaderboards](https://www.tbench.ai/?version=2.1), GPT-5.6 Sol doesn’t appear on the list at all. You can run the benchmark yourself, so OpenAI likely wasn’t making up its numbers, but they’re harder to trust when they aren't public. OpenAI’s graph also shows GPT-5.6 Terra with the same score as Fable 5. But on the leaderboards, Fable 5 is the top performer by resolution rate as of September 1, 2026, scoring 83.8% (which OpenAI’s graph listed as 84.3% on June 26, 2026), and GPT-5.6 Terra places at number six, scoring 78.4% (which OpenAI’s graph also listed as 84.3%).

Once again, OpenAI likely wasn’t lying, as the company could very well have seen the exact numbers it reported when it ran the test, given how AI models change all the time. Furthermore, OpenAI’s graph isn’t particularly specific in the first place, so it’s also possible that the company uses some other metric than the resolution rate on Terminal-Bench’s leaderboards. Regardless, what seems like a straightforward win for GPT-5.6 Sol at a glance becomes a lot less convincing under scrutiny.

## Moving Goalposts: How Version Changes Flip Winners Into Losers

OpenAI discussed GPT-5.6’s performance on Terminal-Bench 2.1, the benchmark's latest version at the time. Since then, Terminal-Bench 3.0 and 4.0 have been released. [Version 3.0 reportedly](https://www.tbench.ai/news/terminal-bench-3-0) “expands on Terminal-Bench 2.1 with more challenging tasks across a wider distribution of valuable work.” And version 4.0 “[has fewer agent timeouts and errors than 3.0, reducing measurement noise.](https://www.tbench.ai/news/terminal-bench-4-0)” 

In June, OpenAI said GPT-5.6 beat Fable 5 on Terminal-Bench 2.1 (91.9% vs 84.3%), but Terminal-Bench’s leaderboards for versions 3.0 and 4.0 tell a different story. [With Terminal-Bench 3.0](https://www.tbench.ai/?version=3.0) as of Sept. 1, GPT-5.6 Sol scores 34.6% for resolution rate, while Fable 5 scores an almost identical 34%. [Terminal-Bench 4.0 further changes the situation:](https://www.tbench.ai/) as of Sept. 1, GPT-5.6 achieves just 37.3% resolution rate, whereas Fable 5 achieves 44.5%.

In just a handful of months (v2.1 released in May, v3.0 released in July, and v4.0 released in August), the same benchmark showed GPT-5.6 beating Fable 5, then performing about the same, and then losing. And keep in mind, as mentioned above, even if this benchmark had a clear, consistent winner, it still wouldn’t necessarily tell you how effectively you can use ChatGPT or Claude to vibe code.

## Grading Their Own Homework: The Inherent Conflict of In-House Tests

Terminal-Bench had “[support with grants](https://www.tbench.ai/contributors)” from a variety of companies, including Anthropic, Google, OpenAI, and more. Maybe that raises an eyebrow, maybe not. But other benchmarks are less independent, making them even harder to trust. 

For example, when OpenAI announced GPT-5.6, it [promised](https://openai.com/index/previewing-gpt-5-6-sol/) “broad improvements in biology workflows” and cited its performance on the GeneBench benchmark over GPT-5.5. Sounds good, but what’s GeneBench? Some kind of academic AI benchmark maintained by researchers and scientists at various top universities? Not quite.

GeneBench is a benchmark [from OpenAI](https://www.biorxiv.org/content/10.64898/2026.04.22.720113v1). Both it and the newer GeneBench Pro have research backing them, but neither has gone through peer review at the time of writing. If I search for GeneBench scores, the results come from third parties rather than official sources, such as [BenchmarkList](https://benchmarklist.com/benchmarks/genebench_pro/), [BenchLM](https://benchlm.ai/benchmarks/genebenchpro), and [LLM](https://llm-stats.com/benchmarks/genebench) Stats. Even then, the only result I found that actually shows the scores of non-OpenAI models has each one performing far worse than OpenAI’s, including older releases like GPT-5.4. So, did OpenAI create a test that favors its models over everybody else’s, or is OpenAI tech just the best for biology workflows?

I certainly can’t tell a difference between prompting GPT-5.6 versus Opus 5 about biology-related topics, but then again, I'm not a biologist. Other companies run different biology-focused tests. For example, when [Anthropic announced Opus 5](https://www.anthropic.com/news/claude-opus-5), it used the BioMysteryBench benchmark to showcase an improvement from 42.4% with Opus 4.8 to 49.4% with Opus 5. Unsurprisingly, [BioMysteryBench comes from Anthropic](https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench).

Conveniently, Anthropic left GPT-5.6’s scores on BioMysteryBench blank in its benchmark scores table. And when OpenAI discussed GPT-5.6’s GeneBench performance, it compared it only with GPT-5.5. In both examples, the companies are advertising products based on tests they create and seemingly don't care to give to their competitors. Even if both of these benchmarks are completely above board and valuable for judging performance across biology workflows, there’s little proof of that right now.

## Higher Scores, Bigger Bills: The Deceptive Economics of Model Upgrades

I’ve written about how [the latest AI models just aren’t worth it](/ai/118992/stop-paying-for-the-newest-ai-models-you-really-dont-need-them-for-most-tasks) for most people and [how deceptive AI pricing is](/ai/119589/i-test-ai-every-day-and-i-still-have-no-idea-what-im-paying-for). AI benchmarks exacerbate both problems. It’s reasonable to think that the benchmark-topping AI models are better than the competition and maybe even worth some extra cash. But all of the aforementioned issues prove otherwise.

For example, [Opus 5 scores a huge 30.2% in ‘Novel problem-solving’](https://www.anthropic.com/news/claude-opus-5) on the ARC-AGI-3 benchmark, while Opus 4.8 scored a paltry 1.5%. Even still, you won’t notice a difference when prompting Opus 4.8 versus Opus 5 for help with your homework. Just looking at the numbers might encourage you to waste your money unnecessarily.

AI usage policies further obfuscate the value proposition of models. For example, [Anthropic provides a variety of benchmark scores](https://www.anthropic.com/news/claude-opus-5) to support its claim that Opus 5 is the company’s “most cost-efficient model” and that it “works more efficiently than other models.” However, a months-long promotion Anthropic has been running (which provides a temporary 50% boost to usage in Claude Code) seriously weakens this claim. So, even if Opus 5 is much more efficient than Anthropic’s other models, you won't be able to get a sense for its true feel. Furthermore, Claude’s five-hour usage limits mean you will often get less done with Opus 5 compared with GPT-5.6, which has only weekly limits on certain plans.

What’s even more frustrating is that you often can’t keep using older models forever, even if they work well for your needs. You also can’t really trust benchmarks to tell you which model you should switch to when your favorite one disappears. So, you either have to pay to try a new model yourself or find someone you trust to run your exact tasks and report back.

## Selective Data and Missing Numbers: The Tricks Hidden in Fine Print

The problems don't end there. The way AI companies display benchmarks is also problematic. I mentioned Opus 4.8’s and Opus 5’s scores in novel problem-solving earlier, but [in that same table](https://www.anthropic.com/news/claude-opus-5), Anthropic left the value for Fable 5’s blank. That's not helpful if you're trying to compare performance.

Other benchmarks have different issues. For example, when [OpenAI announced GPT-5.6](https://openai.com/index/previewing-gpt-5-6-sol/), the company boasted about its cybersecurity chops, calling it “competitive with Mythos Preview using only ~1/3 of the output tokens” on ExploitBench. [ExploitBench](https://exploitbench.ai/) “measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to arbitrary code execution.” 

GPT-5.6's numbers still aren't on ExploitBench’s publicly accessible leaderboard at the time of writing, but toggling between a controlled comparison and comparisons across every test run massively changes the data. The controlled comparison places Mythos Preview at the top with a score of 74%, while GPT-5.5 is far below at just 47%. If you switch to every run, though, Mythos Preview scores 78%, and GPT-5.5 manages 72%.

Chances are good that whatever benchmark an AI company discusses is far from as clear an indication of performance as the company would have you believe.

## Why This Mirage Is Becoming a Real Problem

The way AI companies push deceptive benchmarks is worthy of condemnation. However, the core problem with AI benchmarks is endemic to AI models. You can test a CPU by giving it a tough task, such as video encoding, and measuring how quickly it completes the task. But AI models are all-purpose tools designed to handle any task. As such, what you're measuring is something much closer to intelligence than something as simple as speed.

Furthermore, when most AI models aren’t open-source, who could possibly be better able to test them than their makers? Sure, the general public can and will put AI models through their paces, myself included, but it makes sense that AI companies are uniquely suited to benchmark their technology. Nonetheless, you can't ignore the undeniable conflict of interest between a company creating the tests it uses to sell you on a product's effectiveness.

Just a couple of years ago, the above was more of an existential problem, not an actual one. Back when it was mind-blowing that ChatGPT could write a poem based on a prompt, improvements across each new model version were easy to recognize. Benchmarks then weren’t necessarily any better than they are now, but they weren’t as important.

Today, even models that produce things you can actually see, such as [image](/ai/115006/the-best-ai-image-generators-for-2026) and [video generation](/ai/114227/the-best-ai-video-generators-for-2025) models, aren't drastically different than one another. For example, the upgrade from Nano Banana to [Nano Banana Pro](/ai/114947/nano-banana-pro-unpeeled-see-what-i-made-with-googles-newest-ai-image-generator) was less impressive than the jump to the original Nano Banana. As AI tech improves, you should expect improvements to become increasingly minor, with a greater focus on benchmarks.

## How to Spot the Hype and Protect Your Wallet

Much like a slick car salesman pushing the highest-trimmed model off the lot, AI companies rely on manipulated benchmarks to sell you a fantasy: That the latest release is indispensable, and that you need to pay up right now. But as the numbers show, high scores rarely translate into better real-world workflow. The next time OpenAI or Anthropic drops a shiny chart boasting breakthrough performance, don't buy the hype. Trust your own tests, ignore their self-graded report cards, and always take their benchmark claims with a massive grain of salt.
