cd /news/ai-agents/building-a-rubik-s-cube-decision-are… · home › topics › ai-agents › article
[ARTICLE · art-147398] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Building a Rubik's Cube Decision Arena with 6 Decision Models (Laya, Clef, GLiNER 2.5, Kev and Strands Decider)

A developer built a Rubik's Cube Decision Arena that pits six decision models — Laya, Clef, Clef-Flash, GLiNER 2.5, Kev, and Strands Decider — against the identical scrambled cube and one-move question each turn, running them in parallel through a stdlib HTTP server with per-engine adapters. The project deliberately contrasts decision (system-one) models with the planning/search nature of the cube, enforcing identical conditions and marking any engine that fails to load as unavailable rather than fabricating scores. It reports per-model progress, latency, guided-move and final-state metrics as a scorecard chart.

read15 min views3 publishedOct 8, 2026

Most "AI plays a game" demos hide the interesting part. They show a model winning, and you are left to guess how much of that was the model and how much was the harness. This project does the opposite: it puts six models in a glass box, gives them the exact same scrambled Rubik's cube, the exact same question every turn, and shows you — live, in one screen — exactly where each one succeeds and where it falls apart.

The result is uncomfortable and honest. A Rubik's cube is a planning problem: you need lookahead and search. The models in this arena are decision models: fast, calibrated judgments about the current state. They are not the same thing, and the arena makes that difference visible.

A "decision model" (sometimes called a system-one model) does one thing well: it looks at a situation, weighs a closed set of options, and returns a choice with a confidence. It is fast and it is often surprisingly well-calibrated.

A Rubik's cube is not a decision problem. It is a search problem. From a scrambled state there is an optimal solution length, and finding it means exploring a tree of moves. A model that only ever looks at the current state — with no memory of the plan and no lookahead — is being asked to do something it was not built for.

So the arena asks a precise question:

Given the same cube and the same one-move question, how good are six real decision models at picking moves that actually help?

Everything else in the project exists to answer that question fairly.

The whole comparison is only meaningful if the conditions are identical, so the code enforces two rules everywhere:

unavailable with the exact reason, and the race simply runs without it. There are no placeholder scores and no synthetic results. That second rule sounds obvious until you build a demo. It is very tempting to fill a dead panel with something plausible. The arena refuses.

                          ┌──────────────────────────────────┐
                          │  browser   http://127.0.0.1:8077  │
                          │  ui/index.html  (CSS-3D cubes)    │
                          └───────────────┬──────────────────┘
        GET /            GET /api/engines │ POST /api/run
        GET /api/run/<id>  GET /api/chart/<id>
                                          │
                          ┌───────────────▼──────────────────┐
                          │  server.py   (stdlib HTTP)        │
                          │  Hub: loaded engines + active run │
                          │  CHART_LOCK serialises matplotlib │
                          └───────┬───────────────────┬───────┘
                                  │                   │
             one thread per engine (run in parallel)  │
                                  ▼                   ▼
                    ┌────────────────────┐   ┌─────────────────────────────┐
                    │ arena.run()        │   │ charts.build_scorecard()    │
                    │  guided | greedy   │   │ -> results/chart_<id>.png   │
                    │  | staged          │   │ progress · latency ·        │
                    └─────────┬──────────┘   │ guided · final              │
                              │              └─────────────────────────────┘
              each turn: decide(state, question)
                              ▼
   ┌────────────────────────────────────────────────────────────────┐
   │  engines.py — one contract, six adapters                        │
   │   laya      MLX, in-process          gliner   CPU, in-process   │
   │   strands   HTTP 127.0.0.1:8010      kev      HTTP 127.0.0.1:8008
   │   clef / clef-flash   Cloudflare Workers AI (HTTPS)             │
   └────────────────────────────────────────────────────────────────┘
                              │
                              ▼
   ┌────────────────────────────────────────────────────────────────┐
   │  cube.py  (pycuber state + metrics)                             │
   │  solver.py (kociemba plan / first move)                         │
   └────────────────────────────────────────────────────────────────┘

The stack is deliberately boring: a stdlib HTTP server, a single HTML file, and plain Python. There is no framework, no bundler, and no build step. You can read the whole thing in an afternoon.

Tech stack

Layer Choice
Cube simulation pycuber
Optimal search kociemba (two-phase)
Local runtimes MLX ( laya ), torch/MPS (strands ), transformers/CPU (gliner2 )
Cloud inference Cloudflare Workers AI ( clef ,clef-flash )
HTTP client httpx
Web server Python stdlib http.server +ThreadingHTTPServer
UI single-file HTML + CSS 3D transforms + vanilla JS
Charts matplotlib (Agg )
Terminal rich
Automation Playwright
Env uv + Python 3.12

Everything starts with a deterministic scramble. Given a seed, every engine sees the same cube — and the same cube every time you rerun the demo.

FACES = "URFDLB"
MOVES = [f + s for f in FACES for s in ("", "'", "2")]  # 18 quarter/half turns

def scramble(seed: int, n: int = 25) -> tuple[Cube, list[str]]:
    rng = random.Random(seed)
    seq: list[str] = []
    prev_face = None
    while len(seq) < n:
        mv = rng.choice(MOVES)
        if mv[0] == prev_face:      # avoid trivial same-face repeats
            continue
        prev_face = mv[0]
        seq.append(mv)
    c = Cube()
    for mv in seq:
        c.perform_algo(mv)
    return c, seq

Two metrics are tracked. A sticker is correct when it matches its own face's centre colour (centres never move, so the centre is the solved colour). That gives a smooth 0–54 progress line; complete faces give a chunky 0–6 milestone.

def solved_facelets(cube: Cube) -> int:
    n = 0
    for f in FACES:
        ctr = centre_colour(cube, f)
        face = _face(cube, f)
        for r in range(3):
            for c in range(3):
                if face[r][c].colour == ctr:
                    n += 1
    return n

The model never sees a Cube object. It sees a compact text state, identical for every engine:

def state_text(cube: Cube) -> str:
    rows = facelet_rows(cube)
    body = "\n".join(f"{f}: {rows[f]}" for f in FACES)
    return (
        "3x3 Rubik's cube. Faces U R F D L B, each 9 stickers read row by row "
        "seen from outside. Letters are sticker colours (W Y G B R O). A face is "
        "solved when all 9 stickers match its fixed centre sticker.\n"
        f"{body}\n"
        f"Correct stickers: {solved_facelets(cube)}/54. "
        f"Complete faces: {solved_faces(cube)}/6."
    )

Guided mode needs an optimal-ish solver. The project uses kociemba, but wiring it up has a classic gotcha: pycuber colours its stickers with colour names, and Kociemba wants a 54-character string of face letters in URFDLB order.

Hardcoding the colour→face map is wrong, because pycuber's default cube is not the usual white-up arrangement. The fix is to read the map off the centres, which never move:

FACES = "URFDLB"

def colour_map(cube: Cube) -> dict[str, str]:
    """colour name -> face letter, taken from the fixed centres."""
    return {cube.get_face(f)[1][1].colour: f for f in FACES}

def facelet_string(cube: Cube) -> str:
    """54-char Kociemba facelet string, URFDLB order, row-major from outside."""
    m = colour_map(cube)
    out: list[str] = []
    for f in FACES:
        face = cube.get_face(f)
        for r in range(3):
            for c in range(3):
                out.append(m[face[r][c].colour])
    return "".join(out)

@lru_cache(maxsize=200_000)
def _solve_cached(state: str) -> str:
    return kociemba.solve(state)

Every state is cached by its facelet string. The race revisits states constantly (neutral moves, repeated positions), and Kociemba is by far the slowest part of the loop — the cache is the difference between a snappy demo and a stall.

Two more things worth knowing if you build on this:

kociemba has optimal + a few moves rather than exactly optimal. This is the heart of the project. Six very different runtimes — an in-process MLX model, a torch server, a CPU classifier, two cloud endpoints, and a local OpenAI-ish server — are all reduced to one method:

def decide(self, state: str, question: dict, qid: str = "move") -> Decision:
    ...

@dataclass
class Decision:
    engine: str
    choice: str                      # e.g. "R2"
    confidence: float | None
    probabilities: dict[str, float]  # over the 18 moves (may be empty)
    latency_ms: float
    error: str | None = None

The question is a closed choice over the 18 legal moves, with a description for each — the shape that strands-decider requires, and that the others accept too:

_FACE_NAME = {"U": "up", "R": "right", "F": "front",
              "D": "down", "L": "left", "B": "back"}

MOVE_CRITERIA = {}
for _f, _n in _FACE_NAME.items():
    MOVE_CRITERIA[_f]        = f"turn the {_n} face clockwise 90 degrees"
    MOVE_CRITERIA[_f + "'"]  = f"turn the {_n} face anticlockwise 90 degrees"
    MOVE_CRITERIA[_f + "2"]  = f"turn the {_n} face 180 degrees"

def move_question(instructions: str) -> dict:
    return {"type": "choice", "instructions": instructions,
            "criteria": dict(MOVE_CRITERIA)}

Most engines share an HTTP transport, so adding one is usually a tiny subclass:

class HttpEngine(Engine):
    def probe(self) -> tuple[bool, str | None]:
        """A dead server must read as unavailable, not as 'ready'."""
        if not self.health_url:
            return True, None
        try:
            r = httpx.get(self.health_url, timeout=3.0)
        except Exception as e:
            return False, f"no server at {self.health_url} ({type(e).__name__})"
        return True, None

    def decide(self, state, question, qid="move") -> Decision:
        body = {"state": state, "model": self.model, "questions": {qid: question}}
        t0 = time.perf_counter()
        try:
            r = self.client.post(self.url, json=body)
            r.raise_for_status()
            resp = r.json()
        except Exception as e:
            return Decision(self.name, "", None,
                            latency_ms=(time.perf_counter() - t0) * 1000,
                            error=f"{type(e).__name__}: {e}")
        dt = (time.perf_counter() - t0) * 1000
        choice, conf, probs = _parse_answers(resp, qid)
        return Decision(self.name, choice, conf, probs, dt)

The probe() method is why a dead Kev server shows unavailable: no server at 127.0.0.1:8008 instead of silently reporting "ready" and failing on every move.

Model Params Where How
Laya ( aac6fef/laya-mlx ) 421M local, Apple silicon in-process MLX
Strands Decider 2B 1.9B local, torch/MPS HTTP POST /v1/systemone
GLiNER2.5-Decide 340M local, CPU in-process classify_text
Clef-Flash 9B Cloudflare Workers AI HTTPS
Clef 27B Cloudflare Workers AI HTTPS
Kev-0.8B 0.8B local server :8008 HTTP POST /v1/systemone

Why two run in the cloud: this was built on an Apple M2 mini with 8 GB of RAM. A 27B model does not fit at any quantization, so Clef and Clef-Flash are called on Workers AI where they cost no local RAM. The other four run locally — including on a 16 GB laptop — which is the point of the demo.

The unaided models cannot solve a cube. That is measured, not assumed: after 24 moves each, three different architectures all landed on exactly the sticker count they started from. So guided mode separates the two jobs cleanly — the solver is the search, the model is the decision layer inside it.

plan = solution(cube)                 # a concrete, known-good move list
correct = plan[0]                     # the solver's optimal next move

d = engine.decide(state_text(cube), move_question(GREEDY_Q))
mv = pick_move(d, policy, rng)        # sample the distribution, or take argmax

play, tag, replan = correct, "corrected", False
if mv == correct:
    tag = "hit"                       # matched the solver exactly
elif mv in MOVES:
    probe = cube.copy(); apply_move(probe, mv)
    d_m = len(solution(probe))
    if d_m < len(plan):
        play, tag, replan = mv, "accepted", True      # a different move that helped
    elif d_m == len(plan) and neutral_budget > 0:
        play, tag, replan = mv, "neutral", True       # wasted turn (budgeted)
        neutral_budget -= 1
    else:
        play, tag = correct, "corrected"              # solver overrides
else:
    play, tag = correct, "corrected"                  # unusable answer

apply_move(cube, play)
plan = solution(cube) if replan else plan[1:]

The model gets the same 18-move question as the unaided race, so the two are directly comparable. Each pick is scored:

Outcome Meaning
hit the model picked exactly the solver's next move
accepted a different move that still made strict progress
neutral the move did not change the distance — a wasted turn (budgeted)
corrected the model's move was worse or unusable, so the solver played its own move

The cube is guaranteed to advance because the solver always holds a valid plan, and after the neutral budget is spent every move must strictly reduce the distance.

But here is the subtlety that trips people up: a corrected turn still costs one of the moves you allocated. The solver is doing the solving, and the model is being measured — but the clock runs on every turn, including the ones where the model was overruled. With an optimal distance of 22 and a neutral budget of 6, a guided run can legitimately need up to 28 moves. Set the cap below that and slower models get cut off unsolved while still showing the done status.

Most of these models return a probability for each of the 18 moves. Policy decides what to do with it:

def pick_move(d, policy: str, rng: random.Random) -> str:
    if policy == "sample" and d.probabilities:
        opts = [(m, v) for m, v in d.probabilities.items() if m in MOVES and v > 0]
        if opts:
            total = sum(v for _, v in opts)
            r = rng.random() * total
            acc = 0.0
            for m, v in opts:
                acc += v
                if r <= acc:
                    return m
            return opts[-1][0]
    return d.choice      # argmax

sample`` argmax The server is stdlib-only. Its job is small: hold the loaded engines, run one race at a time, and serve the UI. The interesting part is concurrency.

A run spawns one thread per ready engine, so all six models think in parallel:

def _drive(self, run: dict) -> None:
    threads = []
    for panel in run["panels"]:
        if panel["status"] != "ready":
            continue
        t = threading.Thread(target=self._one_engine, args=(run, panel), daemon=True)
        t.start()
        threads.append(t)
    for t in threads:
        t.join()

Every turn, the engine thread fires an on_move callback that updates the shared panel under a lock. The browser polls /api/run/<id> every 350 ms and redraws the cubes.

There is one nasty bug worth calling out, because it is the kind of thing that only shows up in a demo: matplotlib's pyplot is not thread-safe. The scorecard is built on the run thread when a race finishes, and on demand by the /api/chart/<id> route. If those two ever overlap, the render corrupts and the browser shows a broken image.

The fix is two-fold — a lock, and getting the order right:

CHART_LOCK = threading.Lock()

snapshot = {**run, "status": "done", "finished_at": finished_at}
(RESULTS / f"ui_run_{run['id']}.json").write_text(json.dumps(snapshot, indent=2))
with CHART_LOCK:
    build_scorecard(run, RESULTS / f"chart_{run['id']}.png")
with self.lock:
    run["status"] = "done"          # flipped LAST, only after the PNG exists

The status flips to done after the PNG is on disk, so the page can never request a chart that does not exist yet. The front-end adds a retry with a cache-buster as a belt-and-braces measure.

The API is four routes:

Route Purpose
GET / the single-page UI
GET /api/engines per-engine load status (polled on startup)
POST /api/run start a race {seed, mode, policy, moves, engines}
GET /api/run/<id> current run state (polled while running)
GET /api/chart/<id> the scorecard PNG

The cube is not an image or a canvas — it is six DOM faces positioned in 3D space with CSS transforms, slowly rotating. That is the whole trick:

/* ui/index.html */
.scene { height: 164px; perspective: 1000px; perspective-origin: 50% 45%; }
.cube3d {
  position: relative; width: var(--cube); height: var(--cube);
  transform-style: preserve-3d;
  animation: spin 46s linear infinite;
}
.face3d.U { transform: rotateX(90deg)  translateZ(calc(var(--cube)/2)); }
.face3d.D { transform: rotateX(-90deg) translateZ(calc(var(--cube)/2)); }
.face3d.F { transform: translateZ(calc(var(--cube)/2)); }
.face3d.B { transform: rotateY(180deg) translateZ(calc(var(--cube)/2)); }
.face3d.L { transform: rotateY(-90deg)  translateZ(calc(var(--cube)/2)); }
.face3d.R { transform: rotateY(90deg)   translateZ(calc(var(--cube)/2)); }

Each face is a 3×3 grid of stickers whose colours come straight from the model's net (the per-face colour strings). The layout is three columns by two rows, with a short-screen media query that shrinks the cube so all six stay in one frame.

Two front-end details earned their keep during development:

runId is cleared and the engine poll resumes. The first version re-rendered empty placeholders and startRace() is async, so the status text briefly still reads "models ready". Tests wait on the chart image actually (naturalWidth > 0), not on the status string. When a race finishes, the server renders a four-panel dark scorecard with

matplotlib:

hit / accepted / neutral / corrected bars, so you can see at a glance how much of the solve was the model. Colours are stable per engine key, so a model keeps its colour across runs.

Short version; the full write-up is in REPORT.md.

Unaided (greedy / staged), same scramble, seed 20261002: after 24 moves each, three different architectures landed on exactly the sticker count they started from — zero net progress, and no face ever completed. Even 60 moves of subgoal-named play (staged mode) topped out at 17 of 54 stickers.

Guided mode, same scramble (optimal distance 22): every model solves the cube, because the solver is doing the solving. The honest reading is in the counters:

model solved moves hits accepted neutral solver corrected
laya-mlx yes 28 0 1 6 21
strands-decider yes 28 0 2 6 20
GLiNER2.5-Decide yes 27 4 0 5 18

Read it carefully:

U almost every turn, and the optimal plan happened to start with The headline is not "who won". It is that small decision models do not plan, and the arena shows exactly where that breaks down — while never pretending the model did something it didn't.

A few things that cost real debugging time and are worth stealing:

pip install line, or a failed build aborts the whole command and takes pycuber down with it.pyplot is not thread-safe.probe() is what turns a silent failure into an honest optimal + neutral_budget silently cuts models off unsolved.

git clone https://github.com/harishkotra/cube-arena.git && cd cube-arena
uv venv --python 3.12 && source .venv/bin/activate
uv pip install laya-mlx strands-decider gliner2 pycuber rich matplotlib requests httpx
uv pip install kociemba
python server.py            # open http://127.0.0.1:8077

Everything is optional except pycuber: a model without credentials or a running server simply shows unavailable. ./run.sh starts the strands server and the UI together; ./run_kev.sh starts the Kev server.

There is no framework and no build step, so contributing is easy. The most common contribution is a new engine: subclass HttpEngine (or Engine), register it in ENGINE_SPECS in server.py, and the UI, charts, and CLI pick it up automatically.

Other good first contributions: per-move regret and calibration metrics, a

best-of-N mode, an OpenAI-compatible adapter, a /api/runs history index, and a multi-seed suite. The full list is in the README's "Feature ideas" section.

If you build something on top of this, I'd love to see it. The whole point is that comparison under identical conditions is more interesting than a leaderboard.

Code & more: https://www.dailybuild.xyz/project/277-cube-arena

── more in #ai-agents 4 stories · sorted by recency
── more on @laya 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-a-rubik-s-c…] indexed:0 read:15min 2026-10-08 · —