cd /news/artificial-intelligence/benchmarking-kimi-k3-opus-5-grok-4-5… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-78680] src=quesma.com β†— pub= topic=artificial-intelligence verified=true sentiment=Β· neutral

Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You

Claude Opus 5 achieved 100% pass@3 on the Baba Is You Intro stage, outperforming Kimi K3 (96%), Gemini 3.6 Flash (88%), and Grok 4.5 (75%) in a benchmark by Quesma, while also being more cost-effective than Claude Fable 5. Kimi K3 was cheaper than Opus 5 with nearly the same pass rate, while Gemini 3.6 Flash proved more expensive than its predecessor and Grok 4.5 failed one level.

read6 min views1 publishedJul 29, 2026
Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You
Image: Quesma (auto-discovered)

There are a few exciting model releases: Kimi K3, Grok 4.5, Gemini 3.6 Flash, and Claude Opus 5. It was a fruitful July!

Since they all are reasonably high in the Artificial Analysis Intelligence Index, we re-ran these on our Baba Is Bench, where agents play the lovely puzzle game Baba Is You.

While the game is already 7 years old, trajectories showed no signs of knowing the solution, a sharp contrast with SWE-bench Verified leaks. We do it in two rounds: first, three repetitions for the initial stage, the Intro, using the same Terminus-2 harness. Then we run these on the next stage, the Lake, with custom harnesses.

Stage 0: The Intro #

Let’s start with the simplest one, to have a quick check of speed and cost. We pretty much assumed these models would solve most tasks. For clarity, we dimmed other rows, highlighting the new models.

Model 00 baba is you 01 where do i go? 02 now what is this? 03 out of reach 04 still out of reach 05 volcano 06 off limits 07 grass yard pass@1 ↓ pass@3 ↓ Turns ↓ Output tokens ↓ Total cost ↑
Claude Opus 5 100% 100% 7 23k $17.15
GPT-5.5 100% 100% 8 15k $13.62
Claude Fable 5 100% 100% 6 29k $40.98
GPT-5.6 Sol 100% 100% 11 10k $11.91
Kimi K3 96% 100% 9 29k $12.56
Claude Opus 4.8 96% 100% 35 46k $63.14
GLM-5.2 96% 100% 12 75k $6.20
Gemini 3.1 Pro 92% 100% 8 38k $12.41
Gemini 3.6 Flash 88% 100% 92 82k $124.31
Gemini 3.5 Flash 88% 100% 111 74k $98.31
GPT-5.6 Terra 88% 100% 41 55k $34.28
Grok 4.5 75% 88% 86 68k $50.66
Claude Sonnet 5 67% 88% 22 151k $42.21
GPT-5.6 Luna 63% 88% 118 129k $59.32
Hy3 54% 75% 25 76k $3.23
MiniMax M3 46% 63% 96 173k $34.43
Qwen3.6 27B 29% 38% 24 82k $5.59
DeepSeek V4 Pro 29% 38% 37 183k $11.49
Qwen3.7 Max 25% 38% 21 132k $16.98

Claude Opus 5 not only solved all attempts, but also was more than twice as cheap as Fable 5 (yet, still almost 3x as expensive as GLM-5.2). Kimi K3 was cheaper than Opus, with almost the same pass rate.

Gemini 3.6 Flash… We expected it to be more cost-effective than its ridiculously expensive predecessor, Gemini 3.5 Flash. It was slightly better, yet even more expensive.

Grok 4.5 was neither cheap nor efficient, and even failed at one level.

Stage 1: The Lake #

Per our methodology, we advanced the models with 100% pass@3: Opus 5, Kimi K3, and Gemini 3.6 Flash. As in our Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?, we run all levels just once, to keep costs reasonable.

For each model, we chose their native harness: Claude Code for Opus 5, kimi-cli for Kimi K3, and… well, with Gemini, as previously, there were issues.

Struggling with Gemini

First I tried Antigravity CLI, but, as we already knew, it is incompatible with OpenRouter, which we use to evaluate all other models. Then I (PM) used my Gemini subscription, but it ate it in no time. Then I charged my personal Google account… just to discover that sometimes it is quick (because it web-searches) and other times it goes in loops. Finally I ran it with a standard Terminus-2 harness.

Like its predecessor, Gemini 3.6 Flash is fast (135 tokens per second), but this also means it spends your dollars extremely quickly. On levels it couldn’t solve, it looped without making progress, made thousands of tool calls (1870 in the case of one level) and in one case consumed over 500k tokens of context window.

At peak, the run consumed $7 per minute (extrapolated: over $400 per hour) by aimlessly looping on 9 levels it couldn’t solve. After $260, I (PG) gave this reckless driver a speeding ticket and stopped it before it did more harm to my wallet.

Speed

Measuring against the wall clock, Opus 5 was slower than Fable 5, and Kimi K3 slower than our previous GLM-5.2. One caveat: we tested Kimi K3 right after release, when API quotas were throttling it. Who knows, it may be faster now.

Cost

Cost is where it gets interesting. Opus 5 is between the cheaper GPT-5.6 Sol and the more expensive Fable 5, except for the two hardest tasks, for which it was costlier than Fable 5. So unless you are after the hardest tasks, Opus 5 might be the way to go.

Kimi K3 falls between Opus 4.8 and GPT-5.5, which is rather expensive. So, open weights do not mean free or even cheap: compute to run it is a non-trivial cost.

Gemini 3.6 Flash performs the worst amongst the tested models, and that doesn’t even count the over $260 it spent trying to solve the remainder of the levels. While on this chart it is just $10 of total spend, it does not account for all the loops on unfinished tasks.

Conclusion #

Claude Opus 5 delivers β€” almost as good as Fable 5, yet cheaper. Kimi K3 is clearly a good model, and it is a true miracle that it is now open-weight on Hugging Face.

Grok 4.5 does not generalize: good on official benchmarks, but when doing something slightly different, its score plummets. It is hard to tell whether it is benchmaxxing or just focused on a narrow set of skills.

Gemini fell off. While the releases of Gemini 2.5, 3 and 3.1 were true landmarks, the newest ones do not deliver β€” not in raw skill, pricing, or tooling. We loved Gemini models for thinking, so we hope they will rebound.

On another note, the community needs a better neutral harness for benchmarks than Terminus-2. We also tested Pi and OpenCode, but these were even less efficient. Maybe there is a reason why Frontier-Bench v0.1 (previously called Terminal-Bench 3) lists native harnesses.

And what is your take on the frontier July 2026 models?

Stay tuned for future posts and releases

── more in #artificial-intelligence 4 stories Β· sorted by recency
── more on @kimi k3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/benchmarking-kimi-k3…] indexed:0 read:6min 2026-07-29 Β· β€”