{"slug": "building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2", "title": "Building a Rubik's Cube Decision Arena with 6 Decision Models (Laya, Clef, GLiNER 2.5, Kev and Strands Decider)", "summary": "A developer built a Rubik's Cube Decision Arena that pits six decision models — Laya, Clef, Clef-Flash, GLiNER 2.5, Kev, and Strands Decider — against the identical scrambled cube and one-move question each turn, running them in parallel through a stdlib HTTP server with per-engine adapters. The project deliberately contrasts decision (system-one) models with the planning/search nature of the cube, enforcing identical conditions and marking any engine that fails to load as unavailable rather than fabricating scores. It reports per-model progress, latency, guided-move and final-state metrics as a scorecard chart.", "body_md": "Most \"AI plays a game\" demos hide the interesting part. They show a model winning, and you are left to guess how much of that was the model and how much was the harness. This project does the opposite: it puts six models in a glass box, gives them the **exact same** scrambled Rubik's cube, the **exact same** question every turn, and shows you — live, in one screen — exactly where each one succeeds and where it falls apart.\n\nThe result is uncomfortable and honest. A Rubik's cube is a **planning** problem: you need lookahead and search. The models in this arena are **decision** models: fast, calibrated judgments about the current state. They are not the same thing, and the arena makes that difference visible.\n\nA \"decision model\" (sometimes called a *system-one* model) does one thing well: it looks at a situation, weighs a closed set of options, and returns a choice with a confidence. It is fast and it is often surprisingly well-calibrated.\n\nA Rubik's cube is not a decision problem. It is a search problem. From a scrambled state there is an optimal solution length, and finding it means exploring a tree of moves. A model that only ever looks at the current state — with no memory of the plan and no lookahead — is being asked to do something it was not built for.\n\nSo the arena asks a precise question:\n\nGiven the same cube and the same one-move question, how good are six real decision models at picking moves that actually help?\n\nEverything else in the project exists to answer that question fairly.\n\nThe whole comparison is only meaningful if the conditions are identical, so the code enforces two rules everywhere:\n\n`unavailable` with the exact reason, and the race simply runs without it. There are no placeholder scores and no synthetic results.\nThat second rule sounds obvious until you build a demo. It is very tempting to fill a dead panel with something plausible. The arena refuses.\n\n```\n                          ┌──────────────────────────────────┐\n                          │  browser   http://127.0.0.1:8077  │\n                          │  ui/index.html  (CSS-3D cubes)    │\n                          └───────────────┬──────────────────┘\n        GET /            GET /api/engines │ POST /api/run\n        GET /api/run/<id>  GET /api/chart/<id>\n                                          │\n                          ┌───────────────▼──────────────────┐\n                          │  server.py   (stdlib HTTP)        │\n                          │  Hub: loaded engines + active run │\n                          │  CHART_LOCK serialises matplotlib │\n                          └───────┬───────────────────┬───────┘\n                                  │                   │\n             one thread per engine (run in parallel)  │\n                                  ▼                   ▼\n                    ┌────────────────────┐   ┌─────────────────────────────┐\n                    │ arena.run()        │   │ charts.build_scorecard()    │\n                    │  guided | greedy   │   │ -> results/chart_<id>.png   │\n                    │  | staged          │   │ progress · latency ·        │\n                    └─────────┬──────────┘   │ guided · final              │\n                              │              └─────────────────────────────┘\n              each turn: decide(state, question)\n                              ▼\n   ┌────────────────────────────────────────────────────────────────┐\n   │  engines.py — one contract, six adapters                        │\n   │   laya      MLX, in-process          gliner   CPU, in-process   │\n   │   strands   HTTP 127.0.0.1:8010      kev      HTTP 127.0.0.1:8008\n   │   clef / clef-flash   Cloudflare Workers AI (HTTPS)             │\n   └────────────────────────────────────────────────────────────────┘\n                              │\n                              ▼\n   ┌────────────────────────────────────────────────────────────────┐\n   │  cube.py  (pycuber state + metrics)                             │\n   │  solver.py (kociemba plan / first move)                         │\n   └────────────────────────────────────────────────────────────────┘\n```\n\nThe stack is deliberately boring: a stdlib HTTP server, a single HTML file, and plain Python. There is no framework, no bundler, and no build step. You can read the whole thing in an afternoon.\n\n**Tech stack**\n\n| Layer | Choice | \n|---|---|\n| Cube simulation | `pycuber` | \n| Optimal search | `kociemba` (two-phase) | \n| Local runtimes | MLX ( `laya` ), torch/MPS (`strands` ), transformers/CPU (`gliner2` ) | \n| Cloud inference | Cloudflare Workers AI ( `clef` ,`clef-flash` ) | \n| HTTP client | `httpx` | \n| Web server | Python stdlib `http.server` +`ThreadingHTTPServer` | \n| UI | single-file HTML + CSS 3D transforms + vanilla JS | \n| Charts | `matplotlib` (`Agg` ) | \n| Terminal | `rich` | \n| Automation | Playwright | \n| Env | `uv` + Python 3.12 | \n\nEverything starts with a deterministic scramble. Given a seed, every engine sees the same cube — and the same cube every time you rerun the demo.\n\n```\n# cube.py\nFACES = \"URFDLB\"\nMOVES = [f + s for f in FACES for s in (\"\", \"'\", \"2\")]  # 18 quarter/half turns\n\ndef scramble(seed: int, n: int = 25) -> tuple[Cube, list[str]]:\n    rng = random.Random(seed)\n    seq: list[str] = []\n    prev_face = None\n    while len(seq) < n:\n        mv = rng.choice(MOVES)\n        if mv[0] == prev_face:      # avoid trivial same-face repeats\n            continue\n        prev_face = mv[0]\n        seq.append(mv)\n    c = Cube()\n    for mv in seq:\n        c.perform_algo(mv)\n    return c, seq\n```\n\nTwo metrics are tracked. A **sticker** is correct when it matches its own face's centre colour (centres never move, so the centre *is* the solved colour). That gives a smooth 0–54 progress line; complete faces give a chunky 0–6 milestone.\n\n``` php\n# cube.py\ndef solved_facelets(cube: Cube) -> int:\n    n = 0\n    for f in FACES:\n        ctr = centre_colour(cube, f)\n        face = _face(cube, f)\n        for r in range(3):\n            for c in range(3):\n                if face[r][c].colour == ctr:\n                    n += 1\n    return n\n```\n\nThe model never sees a `Cube` object. It sees a compact text state, identical for every engine:\n\n``` php\n# cube.py\ndef state_text(cube: Cube) -> str:\n    rows = facelet_rows(cube)\n    body = \"\\n\".join(f\"{f}: {rows[f]}\" for f in FACES)\n    return (\n        \"3x3 Rubik's cube. Faces U R F D L B, each 9 stickers read row by row \"\n        \"seen from outside. Letters are sticker colours (W Y G B R O). A face is \"\n        \"solved when all 9 stickers match its fixed centre sticker.\\n\"\n        f\"{body}\\n\"\n        f\"Correct stickers: {solved_facelets(cube)}/54. \"\n        f\"Complete faces: {solved_faces(cube)}/6.\"\n    )\n```\n\nGuided mode needs an optimal-ish solver. The project uses `kociemba`, but wiring it up has a classic gotcha: **pycuber colours its stickers with colour names, and Kociemba wants a 54-character string of *face letters* in URFDLB order.**\n\nHardcoding the colour→face map is wrong, because pycuber's default cube is not the usual white-up arrangement. The fix is to read the map off the centres, which never move:\n\n``` php\n# solver.py\nFACES = \"URFDLB\"\n\ndef colour_map(cube: Cube) -> dict[str, str]:\n    \"\"\"colour name -> face letter, taken from the fixed centres.\"\"\"\n    return {cube.get_face(f)[1][1].colour: f for f in FACES}\n\ndef facelet_string(cube: Cube) -> str:\n    \"\"\"54-char Kociemba facelet string, URFDLB order, row-major from outside.\"\"\"\n    m = colour_map(cube)\n    out: list[str] = []\n    for f in FACES:\n        face = cube.get_face(f)\n        for r in range(3):\n            for c in range(3):\n                out.append(m[face[r][c].colour])\n    return \"\".join(out)\n\n@lru_cache(maxsize=200_000)\ndef _solve_cached(state: str) -> str:\n    return kociemba.solve(state)\n```\n\nEvery state is cached by its facelet string. The race revisits states constantly (neutral moves, repeated positions), and Kociemba is by far the slowest part of the loop — the cache is the difference between a snappy demo and a stall.\n\nTwo more things worth knowing if you build on this:\n\n`kociemba` has `optimal + a few` moves rather than exactly `optimal`.\nThis is the heart of the project. Six very different runtimes — an in-process MLX model, a torch server, a CPU classifier, two cloud endpoints, and a local OpenAI-ish server — are all reduced to **one method**:\n\n``` php\n# engines.py\ndef decide(self, state: str, question: dict, qid: str = \"move\") -> Decision:\n    ...\n\n@dataclass\nclass Decision:\n    engine: str\n    choice: str                      # e.g. \"R2\"\n    confidence: float | None\n    probabilities: dict[str, float]  # over the 18 moves (may be empty)\n    latency_ms: float\n    error: str | None = None\n```\n\nThe question is a closed `choice` over the 18 legal moves, with a description for each — the shape that strands-decider requires, and that the others accept too:\n\n```\n# engines.py\n_FACE_NAME = {\"U\": \"up\", \"R\": \"right\", \"F\": \"front\",\n              \"D\": \"down\", \"L\": \"left\", \"B\": \"back\"}\n\nMOVE_CRITERIA = {}\nfor _f, _n in _FACE_NAME.items():\n    MOVE_CRITERIA[_f]        = f\"turn the {_n} face clockwise 90 degrees\"\n    MOVE_CRITERIA[_f + \"'\"]  = f\"turn the {_n} face anticlockwise 90 degrees\"\n    MOVE_CRITERIA[_f + \"2\"]  = f\"turn the {_n} face 180 degrees\"\n\ndef move_question(instructions: str) -> dict:\n    return {\"type\": \"choice\", \"instructions\": instructions,\n            \"criteria\": dict(MOVE_CRITERIA)}\n```\n\nMost engines share an HTTP transport, so adding one is usually a tiny subclass:\n\n``` php\n# engines.py\nclass HttpEngine(Engine):\n    def probe(self) -> tuple[bool, str | None]:\n        \"\"\"A dead server must read as unavailable, not as 'ready'.\"\"\"\n        if not self.health_url:\n            return True, None\n        try:\n            r = httpx.get(self.health_url, timeout=3.0)\n        except Exception as e:\n            return False, f\"no server at {self.health_url} ({type(e).__name__})\"\n        return True, None\n\n    def decide(self, state, question, qid=\"move\") -> Decision:\n        body = {\"state\": state, \"model\": self.model, \"questions\": {qid: question}}\n        t0 = time.perf_counter()\n        try:\n            r = self.client.post(self.url, json=body)\n            r.raise_for_status()\n            resp = r.json()\n        except Exception as e:\n            return Decision(self.name, \"\", None,\n                            latency_ms=(time.perf_counter() - t0) * 1000,\n                            error=f\"{type(e).__name__}: {e}\")\n        dt = (time.perf_counter() - t0) * 1000\n        choice, conf, probs = _parse_answers(resp, qid)\n        return Decision(self.name, choice, conf, probs, dt)\n```\n\nThe `probe()` method is why a dead Kev server shows `unavailable: no server at 127.0.0.1:8008` instead of silently reporting \"ready\" and failing on every move.\n\n| Model | Params | Where | How | \n|---|---|---|---|\n| Laya ( `aac6fef/laya-mlx` ) | 421M | local, Apple silicon | in-process MLX | \n| Strands Decider 2B | 1.9B | local, torch/MPS | HTTP `POST /v1/systemone` | \n| GLiNER2.5-Decide | 340M | local, CPU | in-process `classify_text` | \n| Clef-Flash | 9B | Cloudflare Workers AI | HTTPS | \n| Clef | 27B | Cloudflare Workers AI | HTTPS | \n| Kev-0.8B | 0.8B | local server `:8008` | HTTP `POST /v1/systemone` | \n\nWhy two run in the cloud: this was built on an Apple M2 mini with 8 GB of RAM. A 27B model does not fit at any quantization, so Clef and Clef-Flash are called on Workers AI where they cost no local RAM. The other four run locally — including on a 16 GB laptop — which is the point of the demo.\n\nThe unaided models cannot solve a cube. That is measured, not assumed: after 24 moves each, three different architectures all landed on exactly the sticker count they started from. So guided mode separates the two jobs cleanly — **the solver is the search, the model is the decision layer inside it.**\n\n```\n# arena.py (abridged)\nplan = solution(cube)                 # a concrete, known-good move list\ncorrect = plan[0]                     # the solver's optimal next move\n\nd = engine.decide(state_text(cube), move_question(GREEDY_Q))\nmv = pick_move(d, policy, rng)        # sample the distribution, or take argmax\n\nplay, tag, replan = correct, \"corrected\", False\nif mv == correct:\n    tag = \"hit\"                       # matched the solver exactly\nelif mv in MOVES:\n    probe = cube.copy(); apply_move(probe, mv)\n    d_m = len(solution(probe))\n    if d_m < len(plan):\n        play, tag, replan = mv, \"accepted\", True      # a different move that helped\n    elif d_m == len(plan) and neutral_budget > 0:\n        play, tag, replan = mv, \"neutral\", True       # wasted turn (budgeted)\n        neutral_budget -= 1\n    else:\n        play, tag = correct, \"corrected\"              # solver overrides\nelse:\n    play, tag = correct, \"corrected\"                  # unusable answer\n\napply_move(cube, play)\nplan = solution(cube) if replan else plan[1:]\n```\n\nThe model gets the **same 18-move question** as the unaided race, so the two are directly comparable. Each pick is scored:\n\n| Outcome | Meaning | \n|---|---|\n| `hit` | the model picked exactly the solver's next move | \n| `accepted` | a different move that still made strict progress | \n| `neutral` | the move did not change the distance — a wasted turn (budgeted) | \n| `corrected` | the model's move was worse or unusable, so the **solver played its own move** | \n\nThe cube is guaranteed to advance because the solver always holds a valid plan, and after the neutral budget is spent every move must strictly reduce the distance.\n\n**But here is the subtlety that trips people up:** a `corrected` turn still costs one of the `moves` you allocated. The solver is doing the solving, and the model is being measured — but the clock runs on every turn, including the ones where the model was overruled. With an optimal distance of 22 and a neutral budget of 6, a guided run can legitimately need up to 28 moves. Set the cap below that and slower models get cut off *unsolved* while still showing the `done` status.\n\nMost of these models return a probability for each of the 18 moves. Policy decides what to do with it:\n\n``` php\n# arena.py\ndef pick_move(d, policy: str, rng: random.Random) -> str:\n    if policy == \"sample\" and d.probabilities:\n        opts = [(m, v) for m, v in d.probabilities.items() if m in MOVES and v > 0]\n        if opts:\n            total = sum(v for _, v in opts)\n            r = rng.random() * total\n            acc = 0.0\n            for m, v in opts:\n                acc += v\n                if r <= acc:\n                    return m\n            return opts[-1][0]\n    return d.choice      # argmax\n```\n\n`sample`` argmax`\nThe server is stdlib-only. Its job is small: hold the loaded engines, run one race at a time, and serve the UI. The interesting part is concurrency.\n\nA run spawns **one thread per ready engine**, so all six models think in parallel:\n\n``` php\n# server.py\ndef _drive(self, run: dict) -> None:\n    threads = []\n    for panel in run[\"panels\"]:\n        if panel[\"status\"] != \"ready\":\n            continue\n        t = threading.Thread(target=self._one_engine, args=(run, panel), daemon=True)\n        t.start()\n        threads.append(t)\n    for t in threads:\n        t.join()\n    # ... persist the run + render the chart, THEN flip status to \"done\"\n```\n\nEvery turn, the engine thread fires an `on_move` callback that updates the shared panel under a lock. The browser polls `/api/run/<id>` every 350 ms and redraws the cubes.\n\nThere is one nasty bug worth calling out, because it is the kind of thing that only shows up in a demo: **matplotlib's `pyplot` is not thread-safe.** The scorecard is built on the run thread when a race finishes, *and* on demand by the `/api/chart/<id>` route. If those two ever overlap, the render corrupts and the browser shows a broken image.\n\nThe fix is two-fold — a lock, and getting the order right:\n\n```\n# server.py\nCHART_LOCK = threading.Lock()\n\n# ... in _drive(), after all engine threads join:\nsnapshot = {**run, \"status\": \"done\", \"finished_at\": finished_at}\n(RESULTS / f\"ui_run_{run['id']}.json\").write_text(json.dumps(snapshot, indent=2))\nwith CHART_LOCK:\n    build_scorecard(run, RESULTS / f\"chart_{run['id']}.png\")\nwith self.lock:\n    run[\"status\"] = \"done\"          # flipped LAST, only after the PNG exists\n```\n\nThe status flips to `done` **after** the PNG is on disk, so the page can never request a chart that does not exist yet. The front-end adds a retry with a cache-buster as a belt-and-braces measure.\n\nThe API is four routes:\n\n| Route | Purpose | \n|---|---|\n| `GET /` | the single-page UI | \n| `GET /api/engines` | per-engine load status (polled on startup) | \n| `POST /api/run` | start a race `{seed, mode, policy, moves, engines}` | \n| `GET /api/run/<id>` | current run state (polled while running) | \n| `GET /api/chart/<id>` | the scorecard PNG | \n\nThe cube is not an image or a canvas — it is six DOM faces positioned in 3D space with CSS transforms, slowly rotating. That is the whole trick:\n\n```\n/* ui/index.html */\n.scene { height: 164px; perspective: 1000px; perspective-origin: 50% 45%; }\n.cube3d {\n  position: relative; width: var(--cube); height: var(--cube);\n  transform-style: preserve-3d;\n  animation: spin 46s linear infinite;\n}\n.face3d.U { transform: rotateX(90deg)  translateZ(calc(var(--cube)/2)); }\n.face3d.D { transform: rotateX(-90deg) translateZ(calc(var(--cube)/2)); }\n.face3d.F { transform: translateZ(calc(var(--cube)/2)); }\n.face3d.B { transform: rotateY(180deg) translateZ(calc(var(--cube)/2)); }\n.face3d.L { transform: rotateY(-90deg)  translateZ(calc(var(--cube)/2)); }\n.face3d.R { transform: rotateY(90deg)   translateZ(calc(var(--cube)/2)); }\n```\n\nEach face is a 3×3 grid of stickers whose colours come straight from the model's `net` (the per-face colour strings). The layout is three columns by two rows, with a short-screen media query that shrinks the cube so all six stay in one frame.\n\nTwo front-end details earned their keep during development:\n\n`runId` is cleared and the engine poll resumes. The first version re-rendered empty placeholders and `startRace()` is async, so the status text briefly still reads \"models ready\". Tests wait on the chart image actually loading (`naturalWidth > 0`), not on the status string.\nWhen a race finishes, the server renders a four-panel dark scorecard with\n\nmatplotlib:\n\n`hit` / `accepted` / `neutral` / `corrected`\nbars, so you can see at a glance how much of the solve was the model.\nColours are stable per engine key, so a model keeps its colour across runs.\n\nShort version; the full write-up is in `REPORT.md`.\n\n**Unaided (greedy / staged), same scramble, seed `20261002`:** after 24 moves each, three different architectures landed on exactly the sticker count they started from — zero net progress, and no face ever completed. Even 60 moves of subgoal-named play (staged mode) topped out at 17 of 54 stickers.\n\n**Guided mode, same scramble (optimal distance 22):** every model solves the cube, because the solver is doing the solving. The honest reading is in the counters:\n\n| model | solved | moves | hits | accepted | neutral | solver corrected | \n|---|---|---|---|---|---|---|\n| laya-mlx | yes | 28 | 0 | 1 | 6 | 21 | \n| strands-decider | yes | 28 | 0 | 2 | 6 | 20 | \n| GLiNER2.5-Decide | yes | 27 | 4 | 0 | 5 | 18 | \n\nRead it carefully:\n\n`U` almost every turn, and the optimal plan happened to start with The headline is not \"who won\". It is that **small decision models do not plan**, and the arena shows exactly where that breaks down — while never pretending the model did something it didn't.\n\nA few things that cost real debugging time and are worth stealing:\n\n`pip install` line, or a failed build aborts the whole command and takes `pycuber` down with it.`pyplot` is not thread-safe.`probe()` is what turns a silent failure into an honest `optimal + neutral_budget` silently cuts models off unsolved.\n\n```\ngit clone https://github.com/harishkotra/cube-arena.git && cd cube-arena\nuv venv --python 3.12 && source .venv/bin/activate\nuv pip install laya-mlx strands-decider gliner2 pycuber rich matplotlib requests httpx\nuv pip install kociemba\npython server.py            # open http://127.0.0.1:8077\n```\n\nEverything is optional except `pycuber`: a model without credentials or a running server simply shows `unavailable`. `./run.sh` starts the strands server and the UI together; `./run_kev.sh` starts the Kev server.\n\nThere is no framework and no build step, so contributing is easy. The most common contribution is a new engine: subclass `HttpEngine` (or `Engine`), register it in `ENGINE_SPECS` in `server.py`, and the UI, charts, and CLI pick it up automatically. \n\nOther good first contributions: per-move regret and calibration metrics, a\n\nbest-of-N mode, an OpenAI-compatible adapter, a `/api/runs` history index, and a multi-seed suite. The full list is in the README's \"Feature ideas\" section.\n\nIf you build something on top of this, I'd love to see it. The whole point is that comparison under identical conditions is more interesting than a leaderboard.\n\nCode & more: [https://www.dailybuild.xyz/project/277-cube-arena](https://www.dailybuild.xyz/project/277-cube-arena)", "url": "https://wpnews.pro/news/building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2", "canonical_source": "https://dev.to/harishkotra/building-a-rubiks-cube-decision-arena-with-6-decision-models-laya-clef-gliner-25-kev-and-2l9i", "published_at": "2026-10-08 07:13:08+00:00", "updated_at": "2026-10-08 07:16:53.137410+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "machine-learning"], "entities": ["Laya", "Clef", "Clef-Flash", "GLiNER 2.5", "Kev", "Strands Decider", "Cloudflare Workers AI", "MLX"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2", "markdown": "https://wpnews.pro/news/building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2.md", "text": "https://wpnews.pro/news/building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2.txt", "jsonld": "https://wpnews.pro/news/building-a-rubik-s-cube-decision-arena-with-6-decision-models-laya-clef-gliner-2.jsonld"}}