# Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You

> Source: <https://quesma.com/blog/baba-kimi-k3-opus-5/>
> Published: 2026-07-29 00:00:00+00:00

# Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You

There are a few exciting model releases: [Kimi K3](https://www.kimi.com/blog/kimi-k3), [Grok 4.5](https://x.ai/news/grok-4-5), [Gemini 3.6 Flash](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/), and [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5). It was a fruitful July!

Since they all are reasonably high in the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/#intelligence), we re-ran these on our [Baba Is Bench](https://quesma.com/blog/baba-is-bench/), where agents play the lovely puzzle game [Baba Is You](https://hempuli.com/baba/).
While the game is already 7 years old, trajectories showed no signs of knowing the solution, a sharp contrast with [SWE-bench Verified leaks](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/).

We do it in two rounds: first, three repetitions for the initial stage, **the Intro**, using the same Terminus-2 harness.
Then we run these on the next stage, **the Lake**, with custom harnesses.

## Stage 0: The Intro

Let’s start with the simplest one, to have a quick check of speed and cost. We pretty much assumed these models would solve most tasks. For clarity, we dimmed other rows, highlighting the new models.

| Model | 00 baba is you | 01 where do i go? | 02 now what is this? | 03 out of reach | 04 still out of reach | 05 volcano | 06 off limits | 07 grass yard | pass@1 ↓ | pass@3 ↓ | Turns ↓ | Output tokens ↓ | Total cost ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 100% | 100% | 7 | 23k | $17.15 | ||||||||
| GPT-5.5 | 100% | 100% | 8 | 15k | $13.62 | ||||||||
| Claude Fable 5 | 100% | 100% | 6 | 29k | $40.98 | ||||||||
| GPT-5.6 Sol | 100% | 100% | 11 | 10k | $11.91 | ||||||||
| Kimi K3 | 96% | 100% | 9 | 29k | $12.56 | ||||||||
| Claude Opus 4.8 | 96% | 100% | 35 | 46k | $63.14 | ||||||||
| GLM-5.2 | 96% | 100% | 12 | 75k | $6.20 | ||||||||
| Gemini 3.1 Pro | 92% | 100% | 8 | 38k | $12.41 | ||||||||
| Gemini 3.6 Flash | 88% | 100% | 92 | 82k | $124.31 | ||||||||
| Gemini 3.5 Flash | 88% | 100% | 111 | 74k | $98.31 | ||||||||
| GPT-5.6 Terra | 88% | 100% | 41 | 55k | $34.28 | ||||||||
| Grok 4.5 | 75% | 88% | 86 | 68k | $50.66 | ||||||||
| Claude Sonnet 5 | 67% | 88% | 22 | 151k | $42.21 | ||||||||
| GPT-5.6 Luna | 63% | 88% | 118 | 129k | $59.32 | ||||||||
| Hy3 | 54% | 75% | 25 | 76k | $3.23 | ||||||||
| MiniMax M3 | 46% | 63% | 96 | 173k | $34.43 | ||||||||
| Qwen3.6 27B | 29% | 38% | 24 | 82k | $5.59 | ||||||||
| DeepSeek V4 Pro | 29% | 38% | 37 | 183k | $11.49 | ||||||||
| Qwen3.7 Max | 25% | 38% | 21 | 132k | $16.98 |

Claude Opus 5 not only solved all attempts, but also was more than twice as cheap as Fable 5 (yet, still almost 3x as expensive as GLM-5.2). Kimi K3 was cheaper than Opus, with almost the same pass rate.

Gemini 3.6 Flash… We expected it to be more cost-effective than its ridiculously expensive predecessor, Gemini 3.5 Flash. It was slightly better, yet even more expensive.

Grok 4.5 was neither cheap nor efficient, and even failed at one level.

## Stage 1: The Lake

Per our methodology, we advanced the models with 100% pass@3: Opus 5, Kimi K3, and Gemini 3.6 Flash.
As in our [Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?](https://quesma.com/blog/baba-is-bench/), we run all levels just once, to keep costs reasonable.

For each model, we chose their native harness: Claude Code for Opus 5, [kimi-cli](https://github.com/MoonshotAI/kimi-cli) for Kimi K3, and… well, with Gemini, as previously, there were issues.

### Struggling with Gemini

First I tried Antigravity CLI, but, as we already knew, it is incompatible with OpenRouter, which we use to evaluate all other models. Then I (PM) used my Gemini subscription, but it ate it in no time. Then I charged my personal Google account… just to discover that sometimes it is quick (because it web-searches) and other times it goes in loops. Finally I ran it with a standard Terminus-2 harness.

Like its predecessor, Gemini 3.6 Flash is fast (135 tokens per second), but this also means it spends your dollars extremely quickly. On levels it couldn’t solve, it looped without making progress, made thousands of tool calls (1870 in the case of one level) and in one case consumed over 500k tokens of context window.

At peak, the run consumed $7 per minute (extrapolated: over $400 per hour) by aimlessly looping on 9 levels it couldn’t solve. After $260, I (PG) gave this reckless driver a speeding ticket and stopped it before it did more harm to my wallet.

### Speed

Measuring against the wall clock, Opus 5 was slower than Fable 5, and Kimi K3 slower than our previous GLM-5.2. One caveat: we tested Kimi K3 right after release, when API quotas were throttling it. Who knows, it may be faster now.

### Cost

Cost is where it gets interesting. Opus 5 is between the cheaper GPT-5.6 Sol and the more expensive Fable 5, except for the two hardest tasks, for which it was costlier than Fable 5. So unless you are after the hardest tasks, Opus 5 might be the way to go.

Kimi K3 falls between Opus 4.8 and GPT-5.5, which is rather expensive. So, open weights do not mean free or even cheap: compute to run it is a non-trivial cost.

Gemini 3.6 Flash performs the worst amongst the tested models, and that doesn’t even count the over $260 it spent trying to solve the remainder of the levels. While on this chart it is just $10 of total spend, it does not account for all the loops on unfinished tasks.

## Conclusion

Claude Opus 5 delivers — almost as good as Fable 5, yet cheaper.
Kimi K3 is clearly a good model, and it is a true miracle that it is [now open-weight on Hugging Face](https://huggingface.co/moonshotai/Kimi-K3).

Grok 4.5 does not generalize: good on official benchmarks, but when doing something slightly different, its score plummets. It is hard to tell whether it is benchmaxxing or just focused on a narrow set of skills.

Gemini fell off. While the releases of Gemini 2.5, 3 and 3.1 were true landmarks, the newest ones do not deliver — not in raw skill, pricing, or tooling. We loved Gemini models for thinking, so we hope they will rebound.

On another note, the community needs a better neutral harness for benchmarks than Terminus-2. We also tested Pi and OpenCode, but these were even less efficient. Maybe there is a reason why [Frontier-Bench v0.1](https://www.frontierbench.ai/) (previously called Terminal-Bench 3) lists native harnesses.

And what is your take on the frontier July 2026 models?

Stay tuned for future posts and releases
