cd /news/artificial-intelligence/kingbench-4-opus-5-5-gpt-6-astra-sol… · home › topics › artificial-intelligence › article
[ARTICLE · art-148745] src=mindstudio.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

KingBench 4: Opus 5.5, GPT-6 Astra, Sol, Haiku 5.5, GLM 5.3 Coding Compared

KingBench 4, an independent benchmark scoring five AI coding models on six difficult software projects out of 10, found Anthropic's Opus 5.5 leading the hardest tasks with wins on the emulator (3.7/10), keyboard app (7.5/10), and room designer (8.5/10). GPT-6.1 Sol produced the strongest Docker-alternative image engine (5.0/10), while GLM 5.3 trailed all rivals and scored 0.6/10 on the emulator after shipping an executable that printed its project name and exited. No model delivered a working DS or 3DS emulator or proved its Docker alternative faster than Docker, and the benchmark used a fixed 60-minute build window per project across 30 total builds.

by read9 min views1 publishedOct 10, 2026
KingBench 4: Opus 5.5, GPT-6 Astra, Sol, Haiku 5.5, GLM 5.3 Coding Compared
Image: Mindstudio (auto-discovered)

KingBench 4 pits five AI coding models against six hard builds, from emulators to a Docker clone. Here's how Opus, Astra, Sol, Haiku, and GLM scored.

What is KingBench 4? #

KingBench 4 is an independent benchmark that tests five AI coding models against six difficult software projects, scoring each result out of 10 based on requested features, real workflows, visual quality, error handling, and performance. The projects include a Game Boy/DS/3DS emulator, a Docker alternative, a 3D room designer, a QMK keyboard configurator, a Pokémon Red recreation, and a fitness app. No model finished every project, but the gaps between them reveal a lot about what “AI writes production code” actually means in practice right now.

TL;DR #

  • Opus 5.5 came out ahead on the hardest tasks, winning the emulator (3.7/10), the keyboard app (7.5/10), and the room designer (8.5/10), while still shipping real bugs like silent save failures and container white-out mishandling.
  • GPT-6.1 Sol built the most credible Docker-alternative image engine (5.0/10) and a solid single-system Game Boy emulator, but repeatedly stopped short of the harder, connected parts of each task.
  • GPT-6 Astra tracked closely behind Sol and Opus across most categories, often delivering polished interfaces that weren’t backed by working hardware or protocol integration.
  • Haiku 5.5 , the smaller Claude model, scored surprisingly close to the larger systems, tying for second on the room designer (8.3/10) and taking second on the keyboard app (6.7/10).
  • GLM 5.3 consistently trailed the other four, including a near-zero emulator score (0.6/10) after shipping an executable that just printed its project name and exited.
  • No model proved its Docker alternative was faster or more efficient than Docker, which was an explicit part of that prompt, and no model delivered a working DS or 3DS emulator despite that being requested alongside Game Boy support.
  • The benchmark used a fixed 60-minute build window per project , fresh sessions for each task, and graded apps “as delivered,” meaning timeouts and incomplete features counted against the final score rather than being patched afterward.

How was KingBench 4 structured? #

Each of the five models built all six projects from the same task prompts and delivery instructions, with a fresh implementation session per project and no carryover between runs. Every run had a 60-minute cutoff, with high reasoning enabled where the coding agent supported it. The models were evaluated through their native coding environments rather than as bare APIs: GPT-6.1 Sol and GPT-6 Astra ran through Codex, Opus 5.5 and Haiku 5.5 ran through Claude Code, and GLM 5.3 ran through Open Code. That makes this a comparison of model-plus-tool combinations, not models in isolation.

The first four models were tested on October 3rd. Haiku 5.5 ran five days later on a newer Claude Code build, which is a small but relevant variable when comparing it directly against the others.

Scoring was out of 10 per project, with all six projects weighted equally toward a final percentage. Crucially, the benchmark creator did not fix missing features before scoring. If a save menu broke, or a keyboard couldn’t actually remap keys, that failure became part of the result. The sample size is small (30 total project builds), which limits how far any single ranking should be generalized across all coding tasks.

Why did the emulator project break every model? #

The emulator prompt asked for a from-scratch Game Boy, DS, and 3DS emulator in Rust, with save states and broad game compatibility. That’s effectively three separate hardware targets bundled into one request, and none of the five models got close to finishing all three.

Sol and Astra both built genuinely functional Game Boy cores, passing 12 selected hardware diagnostics each, with working save states including rejection of corrupted or mismatched data. But neither produced more than file-header inspection for DS or 3DS, so both scored 2.2/10 despite the solid Game Boy layer.

Haiku’s Game Boy implementation passed 10 of 12 diagnostics and cleared 15 serial test cases, but boundary testing exposed timer, interrupt, and sound accuracy failures, plus a malformed save state that crashed on the next rendered frame. Its DS code could execute ARM instructions, but with no surrounding graphics or scheduling system, instruction counters climbed into the millions while video output stayed at zero. It scored 1.8/10.

GLM never reached a usable player at all. Its executable printed the project name and exited, earning it 0.6/10, the lowest score in the category.

Opus 5.5 won at 3.7/10 by actually running two DS homebrew diagnostic programs (Arm Wrestler and Bunny Garden) with working output and deterministic save-state restoration, even though the visuals were rough and 3DS support was absent entirely. The lesson from this category: a well-built Game Boy core with a clean interface doesn’t earn credit for the DS and 3DS work that was also explicitly requested.

Did any model actually build a working Docker alternative? #

Remy is new. The platform isn't. #

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

No. The prompt asked for a tool that could use Docker images, avoid copying Docker’s code or libraries, and beat Docker on efficiency. Every submission fell short of proving the efficiency claim, and independent Linux isolation testing remained incomplete across all five.

Sol scored highest at 5.0/10, correctly handling Docker archive import and export, tag and digest preservation, and scratch image builds, while also passing boundary checks against file-system escape attempts. It lacked published ports, a full interactive terminal, and broad Dockerfile support.

Opus and Astra tied at 4.6/10. Opus implemented the broadest container command set (start, stop, restart, mounts, ports, logs, terminal), but mishandled “white-out” markers, the signals that tell a layer to delete inherited files, in a way that deleted newly created files within the same layer. Astra had a narrower runtime but handled image build, save, load, and tagging workflows cleanly.

GLM and Haiku both scored 3.0/10. GLM could pull and reload a public Alpine image while preserving identity, but its archive extraction allowed writes and deletions outside the intended directory in controlled tests. Haiku’s registry logic worked well, including skipping redundant re-extraction, but a crafted layer could alter permissions and timestamps on files outside the extraction path.

Which model built the best 3D room designer? #

The room designer was the strongest category across the board, with all five models producing editable 3D scenes with furniture placement, custom dimensions, and lighting controls.

Opus won at 8.5/10 with the best material and lighting work in the set: differentiated fabric and wood surfaces, contact shadows, and reflections. It still had boxy furniture and a save-feedback bug where a rejected save was reported as successful even though the stored data didn’t change.

Astra and Haiku tied for second at 8.3/10. Astra’s editor was broad and functional with good import validation, but kept recomputing shadows even when nothing changed, and some walls looked flat. Haiku stood out for render efficiency, pausing drawing when nothing moved and reusing meshes instead of rebuilding them, while also offering strong material variety (wood, upholstery, glass).

Sol scored 8.0/10 with a clean control layout, but a UI bug let a color-swatch button submit the furniture form prematurely, and furniture moves triggered unnecessary full-scene rebuilds.

GLM scored 7.8/10. Its core editor worked, but some wood accents ignored selected colors, foliage looked like stacked geometric shapes, and the import rules disagreed with the UI about minimum room size.

Why did the keyboard configurator expose the biggest gap between looks and function? #

The keyboard project asked for a VIA-style configurator for QMK keyboards with full customization and, critically, the ability to actually communicate with connected firmware. Protocol checks ran against simulated firmware rather than a physical device, so the results confirm protocol-handling logic but not universal hardware compatibility.

Sol and Astra both built attractive, usable interfaces with local editing, undo, and saved profiles, but neither implemented the connected keymap read/write step that is the actual core of a keyboard configurator. Both scored 3.0/10 as a result, a sharp drop from their showing in the room designer.

Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

GLM made a real attempt at device integration, with working layered key reads and writes in testing, but it also rejected standard keyboard definitions, miscoded a macro, and lost some responses until requests timed out, landing at 4.6/10.

Opus won at 7.5/10, implementing device opening, layered keymap reads and writes, standard keyboard definitions, macros, lighting, and backup workflows, verified by reading changed keys back from the simulated device. It still mishandled some malformed replies.

Haiku placed second at 6.7/10 with working key maps, encoders, macros, and backup/restore, but a serious bug let a single unsupported accented character in the macro field crash the entire interface instead of showing a validation message.

Frequently Asked Questions #

What is KingBench 4 testing exactly?

It tests five AI coding models (Opus 5.5, GPT-6.1 Sol, GPT-6 Astra, Haiku 5.5, and GLM 5.3) by having each build the same six complex software projects under a 60-minute time limit, then scoring the delivered apps out of 10 on features, functionality, and error handling.

Which model performed best overall in KingBench 4?

Opus 5.5, run through Claude Code, won the emulator, keyboard configurator, and room designer categories, making it the strongest performer across the projects described. Sol and Astra were close behind on several tasks, and Haiku 5.5 outperformed expectations for a smaller model.

Did any model build a fully working DS or 3DS emulator?

No. All five models were asked for Game Boy, DS, and 3DS support. Several built functional Game Boy cores, but DS support was limited to running a couple of homebrew diagnostic programs at best (Opus), and no model produced working 3DS emulation.

Why did none of the models prove their Docker alternative was more efficient than Docker?

The prompt required a performance and efficiency comparison, but none of the five submissions included benchmark evidence supporting that claim. Each project also had unresolved gaps in container isolation and image layer handling that would need independent verification before any efficiency comparison could be trusted.

Is GLM 5.3 behind the other models in coding ability?

Based on this benchmark, GLM 5.3 scored lowest across most of the reported categories, including a near-failing result on the emulator project. It did make a more substantive attempt at real hardware protocol integration on the keyboard task than some larger models, suggesting its weaknesses are inconsistent rather than uniform.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @kingbench 4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kingbench-4-opus-5-5…] indexed:0 read:9min 2026-10-10 · —