Building a Rubik's Cube Decision Arena with 6 Decision Models (Laya, Clef, GLiNER 2.5, Kev and Strands Decider) A developer built a Rubik's Cube Decision Arena that pits six decision models — Laya, Clef, Clef-Flash, GLiNER 2.5, Kev, and Strands Decider — against the identical scrambled cube and one-move question each turn, running them in parallel through a stdlib HTTP server with per-engine adapters. The project deliberately contrasts decision (system-one) models with the planning/search nature of the cube, enforcing identical conditions and marking any engine that fails to load as unavailable rather than fabricating scores. It reports per-model progress, latency, guided-move and final-state metrics as a scorecard chart. Most "AI plays a game" demos hide the interesting part. They show a model winning, and you are left to guess how much of that was the model and how much was the harness. This project does the opposite: it puts six models in a glass box, gives them the exact same scrambled Rubik's cube, the exact same question every turn, and shows you — live, in one screen — exactly where each one succeeds and where it falls apart. The result is uncomfortable and honest. A Rubik's cube is a planning problem: you need lookahead and search. The models in this arena are decision models: fast, calibrated judgments about the current state. They are not the same thing, and the arena makes that difference visible. A "decision model" sometimes called a system-one model does one thing well: it looks at a situation, weighs a closed set of options, and returns a choice with a confidence. It is fast and it is often surprisingly well-calibrated. A Rubik's cube is not a decision problem. It is a search problem. From a scrambled state there is an optimal solution length, and finding it means exploring a tree of moves. A model that only ever looks at the current state — with no memory of the plan and no lookahead — is being asked to do something it was not built for. So the arena asks a precise question: Given the same cube and the same one-move question, how good are six real decision models at picking moves that actually help? Everything else in the project exists to answer that question fairly. The whole comparison is only meaningful if the conditions are identical, so the code enforces two rules everywhere: unavailable with the exact reason, and the race simply runs without it. There are no placeholder scores and no synthetic results. That second rule sounds obvious until you build a demo. It is very tempting to fill a dead panel with something plausible. The arena refuses. ┌──────────────────────────────────┐ │ browser http://127.0.0.1:8077 │ │ ui/index.html CSS-3D cubes │ └───────────────┬──────────────────┘ GET / GET /api/engines │ POST /api/run GET /api/run/