# Building a Rubik's Cube Decision Arena with 6 Decision Models (Laya, Clef, GLiNER 2.5, Kev and Strands Decider)

> Source: <https://dev.to/harishkotra/building-a-rubiks-cube-decision-arena-with-6-decision-models-laya-clef-gliner-25-kev-and-2l9i>
> Published: 2026-10-08 07:13:08+00:00

Most "AI plays a game" demos hide the interesting part. They show a model winning, and you are left to guess how much of that was the model and how much was the harness. This project does the opposite: it puts six models in a glass box, gives them the **exact same** scrambled Rubik's cube, the **exact same** question every turn, and shows you — live, in one screen — exactly where each one succeeds and where it falls apart.

The result is uncomfortable and honest. A Rubik's cube is a **planning** problem: you need lookahead and search. The models in this arena are **decision** models: fast, calibrated judgments about the current state. They are not the same thing, and the arena makes that difference visible.

A "decision model" (sometimes called a *system-one* model) does one thing well: it looks at a situation, weighs a closed set of options, and returns a choice with a confidence. It is fast and it is often surprisingly well-calibrated.

A Rubik's cube is not a decision problem. It is a search problem. From a scrambled state there is an optimal solution length, and finding it means exploring a tree of moves. A model that only ever looks at the current state — with no memory of the plan and no lookahead — is being asked to do something it was not built for.

So the arena asks a precise question:

Given the same cube and the same one-move question, how good are six real decision models at picking moves that actually help?

Everything else in the project exists to answer that question fairly.

The whole comparison is only meaningful if the conditions are identical, so the code enforces two rules everywhere:

`unavailable` with the exact reason, and the race simply runs without it. There are no placeholder scores and no synthetic results.
That second rule sounds obvious until you build a demo. It is very tempting to fill a dead panel with something plausible. The arena refuses.

```
                          ┌──────────────────────────────────┐
                          │  browser   http://127.0.0.1:8077  │
                          │  ui/index.html  (CSS-3D cubes)    │
                          └───────────────┬──────────────────┘
        GET /            GET /api/engines │ POST /api/run
        GET /api/run/<id>  GET /api/chart/<id>
                                          │
                          ┌───────────────▼──────────────────┐
                          │  server.py   (stdlib HTTP)        │
                          │  Hub: loaded engines + active run │
                          │  CHART_LOCK serialises matplotlib │
                          └───────┬───────────────────┬───────┘
                                  │                   │
             one thread per engine (run in parallel)  │
                                  ▼                   ▼
                    ┌────────────────────┐   ┌─────────────────────────────┐
                    │ arena.run()        │   │ charts.build_scorecard()    │
                    │  guided | greedy   │   │ -> results/chart_<id>.png   │
                    │  | staged          │   │ progress · latency ·        │
                    └─────────┬──────────┘   │ guided · final              │
                              │              └─────────────────────────────┘
              each turn: decide(state, question)
                              ▼
   ┌────────────────────────────────────────────────────────────────┐
   │  engines.py — one contract, six adapters                        │
   │   laya      MLX, in-process          gliner   CPU, in-process   │
   │   strands   HTTP 127.0.0.1:8010      kev      HTTP 127.0.0.1:8008
   │   clef / clef-flash   Cloudflare Workers AI (HTTPS)             │
   └────────────────────────────────────────────────────────────────┘
                              │
                              ▼
   ┌────────────────────────────────────────────────────────────────┐
   │  cube.py  (pycuber state + metrics)                             │
   │  solver.py (kociemba plan / first move)                         │
   └────────────────────────────────────────────────────────────────┘
```

The stack is deliberately boring: a stdlib HTTP server, a single HTML file, and plain Python. There is no framework, no bundler, and no build step. You can read the whole thing in an afternoon.

**Tech stack**

| Layer | Choice | 
|---|---|
| Cube simulation | `pycuber` | 
| Optimal search | `kociemba` (two-phase) | 
| Local runtimes | MLX ( `laya` ), torch/MPS (`strands` ), transformers/CPU (`gliner2` ) | 
| Cloud inference | Cloudflare Workers AI ( `clef` ,`clef-flash` ) | 
| HTTP client | `httpx` | 
| Web server | Python stdlib `http.server` +`ThreadingHTTPServer` | 
| UI | single-file HTML + CSS 3D transforms + vanilla JS | 
| Charts | `matplotlib` (`Agg` ) | 
| Terminal | `rich` | 
| Automation | Playwright | 
| Env | `uv` + Python 3.12 | 

Everything starts with a deterministic scramble. Given a seed, every engine sees the same cube — and the same cube every time you rerun the demo.

```
# cube.py
FACES = "URFDLB"
MOVES = [f + s for f in FACES for s in ("", "'", "2")]  # 18 quarter/half turns

def scramble(seed: int, n: int = 25) -> tuple[Cube, list[str]]:
    rng = random.Random(seed)
    seq: list[str] = []
    prev_face = None
    while len(seq) < n:
        mv = rng.choice(MOVES)
        if mv[0] == prev_face:      # avoid trivial same-face repeats
            continue
        prev_face = mv[0]
        seq.append(mv)
    c = Cube()
    for mv in seq:
        c.perform_algo(mv)
    return c, seq
```

Two metrics are tracked. A **sticker** is correct when it matches its own face's centre colour (centres never move, so the centre *is* the solved colour). That gives a smooth 0–54 progress line; complete faces give a chunky 0–6 milestone.

``` php
# cube.py
def solved_facelets(cube: Cube) -> int:
    n = 0
    for f in FACES:
        ctr = centre_colour(cube, f)
        face = _face(cube, f)
        for r in range(3):
            for c in range(3):
                if face[r][c].colour == ctr:
                    n += 1
    return n
```

The model never sees a `Cube` object. It sees a compact text state, identical for every engine:

``` php
# cube.py
def state_text(cube: Cube) -> str:
    rows = facelet_rows(cube)
    body = "\n".join(f"{f}: {rows[f]}" for f in FACES)
    return (
        "3x3 Rubik's cube. Faces U R F D L B, each 9 stickers read row by row "
        "seen from outside. Letters are sticker colours (W Y G B R O). A face is "
        "solved when all 9 stickers match its fixed centre sticker.\n"
        f"{body}\n"
        f"Correct stickers: {solved_facelets(cube)}/54. "
        f"Complete faces: {solved_faces(cube)}/6."
    )
```

Guided mode needs an optimal-ish solver. The project uses `kociemba`, but wiring it up has a classic gotcha: **pycuber colours its stickers with colour names, and Kociemba wants a 54-character string of *face letters* in URFDLB order.**

Hardcoding the colour→face map is wrong, because pycuber's default cube is not the usual white-up arrangement. The fix is to read the map off the centres, which never move:

``` php
# solver.py
FACES = "URFDLB"

def colour_map(cube: Cube) -> dict[str, str]:
    """colour name -> face letter, taken from the fixed centres."""
    return {cube.get_face(f)[1][1].colour: f for f in FACES}

def facelet_string(cube: Cube) -> str:
    """54-char Kociemba facelet string, URFDLB order, row-major from outside."""
    m = colour_map(cube)
    out: list[str] = []
    for f in FACES:
        face = cube.get_face(f)
        for r in range(3):
            for c in range(3):
                out.append(m[face[r][c].colour])
    return "".join(out)

@lru_cache(maxsize=200_000)
def _solve_cached(state: str) -> str:
    return kociemba.solve(state)
```

Every state is cached by its facelet string. The race revisits states constantly (neutral moves, repeated positions), and Kociemba is by far the slowest part of the loop — the cache is the difference between a snappy demo and a stall.

Two more things worth knowing if you build on this:

`kociemba` has `optimal + a few` moves rather than exactly `optimal`.
This is the heart of the project. Six very different runtimes — an in-process MLX model, a torch server, a CPU classifier, two cloud endpoints, and a local OpenAI-ish server — are all reduced to **one method**:

``` php
# engines.py
def decide(self, state: str, question: dict, qid: str = "move") -> Decision:
    ...

@dataclass
class Decision:
    engine: str
    choice: str                      # e.g. "R2"
    confidence: float | None
    probabilities: dict[str, float]  # over the 18 moves (may be empty)
    latency_ms: float
    error: str | None = None
```

The question is a closed `choice` over the 18 legal moves, with a description for each — the shape that strands-decider requires, and that the others accept too:

```
# engines.py
_FACE_NAME = {"U": "up", "R": "right", "F": "front",
              "D": "down", "L": "left", "B": "back"}

MOVE_CRITERIA = {}
for _f, _n in _FACE_NAME.items():
    MOVE_CRITERIA[_f]        = f"turn the {_n} face clockwise 90 degrees"
    MOVE_CRITERIA[_f + "'"]  = f"turn the {_n} face anticlockwise 90 degrees"
    MOVE_CRITERIA[_f + "2"]  = f"turn the {_n} face 180 degrees"

def move_question(instructions: str) -> dict:
    return {"type": "choice", "instructions": instructions,
            "criteria": dict(MOVE_CRITERIA)}
```

Most engines share an HTTP transport, so adding one is usually a tiny subclass:

``` php
# engines.py
class HttpEngine(Engine):
    def probe(self) -> tuple[bool, str | None]:
        """A dead server must read as unavailable, not as 'ready'."""
        if not self.health_url:
            return True, None
        try:
            r = httpx.get(self.health_url, timeout=3.0)
        except Exception as e:
            return False, f"no server at {self.health_url} ({type(e).__name__})"
        return True, None

    def decide(self, state, question, qid="move") -> Decision:
        body = {"state": state, "model": self.model, "questions": {qid: question}}
        t0 = time.perf_counter()
        try:
            r = self.client.post(self.url, json=body)
            r.raise_for_status()
            resp = r.json()
        except Exception as e:
            return Decision(self.name, "", None,
                            latency_ms=(time.perf_counter() - t0) * 1000,
                            error=f"{type(e).__name__}: {e}")
        dt = (time.perf_counter() - t0) * 1000
        choice, conf, probs = _parse_answers(resp, qid)
        return Decision(self.name, choice, conf, probs, dt)
```

The `probe()` method is why a dead Kev server shows `unavailable: no server at 127.0.0.1:8008` instead of silently reporting "ready" and failing on every move.

| Model | Params | Where | How | 
|---|---|---|---|
| Laya ( `aac6fef/laya-mlx` ) | 421M | local, Apple silicon | in-process MLX | 
| Strands Decider 2B | 1.9B | local, torch/MPS | HTTP `POST /v1/systemone` | 
| GLiNER2.5-Decide | 340M | local, CPU | in-process `classify_text` | 
| Clef-Flash | 9B | Cloudflare Workers AI | HTTPS | 
| Clef | 27B | Cloudflare Workers AI | HTTPS | 
| Kev-0.8B | 0.8B | local server `:8008` | HTTP `POST /v1/systemone` | 

Why two run in the cloud: this was built on an Apple M2 mini with 8 GB of RAM. A 27B model does not fit at any quantization, so Clef and Clef-Flash are called on Workers AI where they cost no local RAM. The other four run locally — including on a 16 GB laptop — which is the point of the demo.

The unaided models cannot solve a cube. That is measured, not assumed: after 24 moves each, three different architectures all landed on exactly the sticker count they started from. So guided mode separates the two jobs cleanly — **the solver is the search, the model is the decision layer inside it.**

```
# arena.py (abridged)
plan = solution(cube)                 # a concrete, known-good move list
correct = plan[0]                     # the solver's optimal next move

d = engine.decide(state_text(cube), move_question(GREEDY_Q))
mv = pick_move(d, policy, rng)        # sample the distribution, or take argmax

play, tag, replan = correct, "corrected", False
if mv == correct:
    tag = "hit"                       # matched the solver exactly
elif mv in MOVES:
    probe = cube.copy(); apply_move(probe, mv)
    d_m = len(solution(probe))
    if d_m < len(plan):
        play, tag, replan = mv, "accepted", True      # a different move that helped
    elif d_m == len(plan) and neutral_budget > 0:
        play, tag, replan = mv, "neutral", True       # wasted turn (budgeted)
        neutral_budget -= 1
    else:
        play, tag = correct, "corrected"              # solver overrides
else:
    play, tag = correct, "corrected"                  # unusable answer

apply_move(cube, play)
plan = solution(cube) if replan else plan[1:]
```

The model gets the **same 18-move question** as the unaided race, so the two are directly comparable. Each pick is scored:

| Outcome | Meaning | 
|---|---|
| `hit` | the model picked exactly the solver's next move | 
| `accepted` | a different move that still made strict progress | 
| `neutral` | the move did not change the distance — a wasted turn (budgeted) | 
| `corrected` | the model's move was worse or unusable, so the **solver played its own move** | 

The cube is guaranteed to advance because the solver always holds a valid plan, and after the neutral budget is spent every move must strictly reduce the distance.

**But here is the subtlety that trips people up:** a `corrected` turn still costs one of the `moves` you allocated. The solver is doing the solving, and the model is being measured — but the clock runs on every turn, including the ones where the model was overruled. With an optimal distance of 22 and a neutral budget of 6, a guided run can legitimately need up to 28 moves. Set the cap below that and slower models get cut off *unsolved* while still showing the `done` status.

Most of these models return a probability for each of the 18 moves. Policy decides what to do with it:

``` php
# arena.py
def pick_move(d, policy: str, rng: random.Random) -> str:
    if policy == "sample" and d.probabilities:
        opts = [(m, v) for m, v in d.probabilities.items() if m in MOVES and v > 0]
        if opts:
            total = sum(v for _, v in opts)
            r = rng.random() * total
            acc = 0.0
            for m, v in opts:
                acc += v
                if r <= acc:
                    return m
            return opts[-1][0]
    return d.choice      # argmax
```

`sample`` argmax`
The server is stdlib-only. Its job is small: hold the loaded engines, run one race at a time, and serve the UI. The interesting part is concurrency.

A run spawns **one thread per ready engine**, so all six models think in parallel:

``` php
# server.py
def _drive(self, run: dict) -> None:
    threads = []
    for panel in run["panels"]:
        if panel["status"] != "ready":
            continue
        t = threading.Thread(target=self._one_engine, args=(run, panel), daemon=True)
        t.start()
        threads.append(t)
    for t in threads:
        t.join()
    # ... persist the run + render the chart, THEN flip status to "done"
```

Every turn, the engine thread fires an `on_move` callback that updates the shared panel under a lock. The browser polls `/api/run/<id>` every 350 ms and redraws the cubes.

There is one nasty bug worth calling out, because it is the kind of thing that only shows up in a demo: **matplotlib's `pyplot` is not thread-safe.** The scorecard is built on the run thread when a race finishes, *and* on demand by the `/api/chart/<id>` route. If those two ever overlap, the render corrupts and the browser shows a broken image.

The fix is two-fold — a lock, and getting the order right:

```
# server.py
CHART_LOCK = threading.Lock()

# ... in _drive(), after all engine threads join:
snapshot = {**run, "status": "done", "finished_at": finished_at}
(RESULTS / f"ui_run_{run['id']}.json").write_text(json.dumps(snapshot, indent=2))
with CHART_LOCK:
    build_scorecard(run, RESULTS / f"chart_{run['id']}.png")
with self.lock:
    run["status"] = "done"          # flipped LAST, only after the PNG exists
```

The status flips to `done` **after** the PNG is on disk, so the page can never request a chart that does not exist yet. The front-end adds a retry with a cache-buster as a belt-and-braces measure.

The API is four routes:

| Route | Purpose | 
|---|---|
| `GET /` | the single-page UI | 
| `GET /api/engines` | per-engine load status (polled on startup) | 
| `POST /api/run` | start a race `{seed, mode, policy, moves, engines}` | 
| `GET /api/run/<id>` | current run state (polled while running) | 
| `GET /api/chart/<id>` | the scorecard PNG | 

The cube is not an image or a canvas — it is six DOM faces positioned in 3D space with CSS transforms, slowly rotating. That is the whole trick:

```
/* ui/index.html */
.scene { height: 164px; perspective: 1000px; perspective-origin: 50% 45%; }
.cube3d {
  position: relative; width: var(--cube); height: var(--cube);
  transform-style: preserve-3d;
  animation: spin 46s linear infinite;
}
.face3d.U { transform: rotateX(90deg)  translateZ(calc(var(--cube)/2)); }
.face3d.D { transform: rotateX(-90deg) translateZ(calc(var(--cube)/2)); }
.face3d.F { transform: translateZ(calc(var(--cube)/2)); }
.face3d.B { transform: rotateY(180deg) translateZ(calc(var(--cube)/2)); }
.face3d.L { transform: rotateY(-90deg)  translateZ(calc(var(--cube)/2)); }
.face3d.R { transform: rotateY(90deg)   translateZ(calc(var(--cube)/2)); }
```

Each face is a 3×3 grid of stickers whose colours come straight from the model's `net` (the per-face colour strings). The layout is three columns by two rows, with a short-screen media query that shrinks the cube so all six stay in one frame.

Two front-end details earned their keep during development:

`runId` is cleared and the engine poll resumes. The first version re-rendered empty placeholders and `startRace()` is async, so the status text briefly still reads "models ready". Tests wait on the chart image actually loading (`naturalWidth > 0`), not on the status string.
When a race finishes, the server renders a four-panel dark scorecard with

matplotlib:

`hit` / `accepted` / `neutral` / `corrected`
bars, so you can see at a glance how much of the solve was the model.
Colours are stable per engine key, so a model keeps its colour across runs.

Short version; the full write-up is in `REPORT.md`.

**Unaided (greedy / staged), same scramble, seed `20261002`:** after 24 moves each, three different architectures landed on exactly the sticker count they started from — zero net progress, and no face ever completed. Even 60 moves of subgoal-named play (staged mode) topped out at 17 of 54 stickers.

**Guided mode, same scramble (optimal distance 22):** every model solves the cube, because the solver is doing the solving. The honest reading is in the counters:

| model | solved | moves | hits | accepted | neutral | solver corrected | 
|---|---|---|---|---|---|---|
| laya-mlx | yes | 28 | 0 | 1 | 6 | 21 | 
| strands-decider | yes | 28 | 0 | 2 | 6 | 20 | 
| GLiNER2.5-Decide | yes | 27 | 4 | 0 | 5 | 18 | 

Read it carefully:

`U` almost every turn, and the optimal plan happened to start with The headline is not "who won". It is that **small decision models do not plan**, and the arena shows exactly where that breaks down — while never pretending the model did something it didn't.

A few things that cost real debugging time and are worth stealing:

`pip install` line, or a failed build aborts the whole command and takes `pycuber` down with it.`pyplot` is not thread-safe.`probe()` is what turns a silent failure into an honest `optimal + neutral_budget` silently cuts models off unsolved.

```
git clone https://github.com/harishkotra/cube-arena.git && cd cube-arena
uv venv --python 3.12 && source .venv/bin/activate
uv pip install laya-mlx strands-decider gliner2 pycuber rich matplotlib requests httpx
uv pip install kociemba
python server.py            # open http://127.0.0.1:8077
```

Everything is optional except `pycuber`: a model without credentials or a running server simply shows `unavailable`. `./run.sh` starts the strands server and the UI together; `./run_kev.sh` starts the Kev server.

There is no framework and no build step, so contributing is easy. The most common contribution is a new engine: subclass `HttpEngine` (or `Engine`), register it in `ENGINE_SPECS` in `server.py`, and the UI, charts, and CLI pick it up automatically. 

Other good first contributions: per-move regret and calibration metrics, a

best-of-N mode, an OpenAI-compatible adapter, a `/api/runs` history index, and a multi-seed suite. The full list is in the README's "Feature ideas" section.

If you build something on top of this, I'd love to see it. The whole point is that comparison under identical conditions is more interesting than a leaderboard.

Code & more: [https://www.dailybuild.xyz/project/277-cube-arena](https://www.dailybuild.xyz/project/277-cube-arena)
