cd /news/artificial-intelligence/godot-benchmark-2-opus-5-sol-terra · home topics artificial-intelligence article
[ARTICLE · art-74719] src=ziva.sh ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Godot Benchmark 2: Opus 5 > Sol > Terra

Ziva's Godot Benchmark 2 finds Claude Opus 5 is the only model that produces a playable 3D vampire survivor game, scoring 8/10 for world quality and 9/10 for conversation, but costing $65.64 and taking 56 minutes — 7x more expensive and 2x longer than GPT 5.6 Sol. GPT 5.6 Sol ($9.79, 28 min) is playable but rough, while GPT 5.6 Terra ($2.27, 11 min) is unusable because attacks never connect. Ziva recommends Sol for developers actively in the game development loop.

read2 min views1 publishedJul 26, 2026
Godot Benchmark 2: Opus 5 > Sol > Terra
Image: source

Round 2 of our Godot benchmark: Claude Opus 5, GPT 5.6 Sol, and GPT 5.6 Terra each got one prompt to build a 3D vampire survivor in Godot from a full AI-generated game spec, using the bundled KayKit asset packs. All at high reasoning.

TL;DR #

Opus 5 is the only one that feels like a game, but it cost 7x more and took 2x longer than Sol. Sol is playable but rough. Terra is unusable: you can’t hit anything, so you can’t progress.

Model World Conversation Cost Time
Claude Opus 5 8 9 $65.64 56 min
GPT 5.6 Sol 5 6 $9.79 28 min
GPT 5.6 Terra 2 3 $2.27 11 min

The Prompt #

Create me a 3d vampire survivor game. The game spec is in the project root.

The project root held a detailed game spec generated by ChatGPT .

Created Worlds #

Claude Opus 5 GPT 5.6 Sol GPT 5.6 Terra

PlayPlayAll three used the KayKit packs. Opus 5 gives you a dense graveyard forest with a real swarm to kite. Sol works but is rough: the UI isn’t centred, left click does nothing despite the game saying it should, there are no damage indicators, and nothing animates when you move. Terra you can’t play at all — the attacks never connect, so you die without being able to do anything.

Conversation Analysis #

Fable 5 graded each transcript on whether the model worked like a real game dev: scenes over runtime-generated everything, real reasoning about 3D placement, playtests it actually looked at.

Claude Opus 5— 16 scenes, decorations seed-scattered with a combat-zone exclusion, and it looked at 11 of 16 playtests — even catching a 1-FPS collapse from a frame. Docked for building all UI in code.9/10·conversation** GPT 5.6 Sol**— 7 real scenes, but props placed on two perfect rings (hence the empty arena), 45% of the code in one file, and only 1 of 12 playtests reacted to what was on screen — so it shipped bugs it believed it had fixed.6/10·conversation** GPT 5.6 Terra**— one scene, one 807-linemain.gd

holding the whole game, props on a grid formula, and none of its 6 playtests engaged with the output — which is how you ship attacks that never connect.3/10·conversation

My Thoughts #

Opus 5’s is the only one I’d call a game - it was genuinely pretty fun to play. But an hour and $65 per iteration isn’t great. Sol gets you something playable for $10 much quicker, and I would personally recommend using that over Opus if you’re actively part of the game development loop; something we always encourage at Ziva.

Credits #

Assets by KayKit (Adventurers, Forest Nature packs). The game spec was generated with one prompt to ChatGPT: conversation

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ziva 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/godot-benchmark-2-op…] indexed:0 read:2min 2026-07-26 ·