{"slug": "how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks", "title": "How I built a virtual chess coach: Stockfish calculates, Claude explains, ElevenLabs speaks", "summary": "A developer built a virtual chess coach that pairs Stockfish for move calculation with Claude for natural-language explanations and ElevenLabs for spoken output in 22 languages. The system feeds the LLM pre-computed engine facts rather than raw FEN positions and validates every move, line, and evaluation against the board, regenerating answers once when validation fails. The coach runs Stockfish both in-browser via WASM and as a server-side pool of three native processes with a 3,000-position LRU cache.", "body_md": "Every time I lose a game and open the engine analysis, I get the same thing: **−2.3** and an arrow. Fine, my move was bad. But *why*? What should I have seen? Stockfish knows the answer better than any grandmaster — it just can't say it in words.\n\nSo I built a coach that sees the board like Stockfish, explains like a human, and talks to you on a video-call-style screen. The first version came together quickly. Teaching it not to lie took much longer.\n\nThis post covers the architecture, how I feed facts to the LLM instead of positions, how every answer gets validated against the board, voice in 22 languages, costs, and the bugs. There were more bugs than code.\n\nLLMs are bad at chess. They lose track of the board, move pieces that aren't there, and confidently evaluate positions they never calculated. For a coach, that's worse than silence: a 1300 player won't notice that \"the knight on f3 is defended by the bishop\" is false, and will remember the wrong idea.\n\nSo the one rule that never changed:\n\n**Every move, line, and evaluation in an answer comes from the engine. The model only explains.**\n\n```\nBrowser                                   Server (Java / Spring)\n───────                                   ──────────────────────\nStockfish WASM (Web Worker)               Pool of native Stockfish (3 processes)\n ├─ coach's own moves (MultiPV 6)          ├─ game review (depth 14, recheck at 18)\n └─ judging the player's move (depth 12)   └─ answering questions (depth 14, top 3 moves)\n                                                    │  facts (JSON)\nQuestion (text / voice) ─────────────►  Claude Sonnet 5 + \"analyze\" tool\n                                                    │  text\n                                         Answer validation ──► one regeneration\n                                                    │\nAvatar + audio ◄──────────── mp3 ◄────  ElevenLabs v4 Turbo (+ R2 cache)\n```\n\nStockfish runs in two places. In the **browser** (WASM in a Web Worker) for anything that must feel instant and cost nothing: the coach's moves when you play against it, and judging your move during the game. On the **server** as a pool of native processes. The first version spawned one Stockfish per question — half a second to start, which seemed fine until ten questions arrived at once; the small instance ran out of memory, and the whole site went down. Now it's a pool of 3 (64 MB hash each), a 20-second queue timeout, and an LRU cache of 3,000 positions, because the same opening positions get asked about constantly.\n\nMy first version sent the FEN and asked: \"explain why this move is bad\". The model stared at `r1bqkb1r/pppp1ppp/2n2n2/...` and made things up. Once it said 6...Nd4 \"attacks the knight on f3 and the bishop on e2\". There was a queen on e2.\n\nNow the model gets pre-computed facts:\n\n```\n{\n  \"current\": {\n    \"move\": \"14...Bxe4\",\n    \"eval_before\": \"about equal\",\n    \"eval_after\": \"White is winning (about +2.5, a piece up)\",\n    \"engine_best\": \"14...Re8\",\n    \"engine_reason\": \"after Bxe4 the reply Nxe4 wins the bishop: it has no retreat\",\n    \"reply_move\": \"15.Nxe4\",\n    \"engine_top_before\": [\"Re8\", \"h6\", \"Qd7\"]\n  },\n  \"habit_tags\": [\"hanging_piece\"]\n}\n```\n\nTwo non-obvious lessons:\n\n`+0.2` the model sometimes flipped it to the other side. My favorite line from the logs: When the player asks \"what if I take with the rook?\", the model has a tool:\n\n```\n{\n  \"name\": \"analyze\",\n  \"input_schema\": {\n    \"properties\": {\n      \"start\": { \"enum\": [\"before\", \"after\"] },\n      \"moves\": { \"type\": \"array\", \"items\": { \"type\": \"string\" } },\n      \"label\": { \"type\": \"string\" }\n    }\n  }\n}\n```\n\nThe server plays the moves on a real board (max 12 plies), runs Stockfish, and returns the eval, a 6-ply best line, and what the last move attacks. An illegal move returns an error with a note: *\"Do not guess why; say only that the move is not possible there.\"* That note exists because in live games `analyze` once defaulted to the position *before* the player's move, so a perfectly legal move came back \"illegal\" — and the coach confidently explained: *\"c3 is taken by your knight.\"*\n\nEven with facts and a tool, the model sometimes named moves that existed nowhere. So every answer is checked before the player sees it. All moves that appeared anywhere (the game, engine lines, `analyze` results) go into an `allowed` set, and every move is pulled out of the answer:\n\n```\n// Russian piece letters too (Кр, Ф, Л, С, К): the model slipped a whole invented line past the\n// check once by writing \"Фxg5\".\nprivate static final Pattern MOVE = Pattern.compile(\n    \"(?<![A-Za-z0-9А-Яа-я])(O-O-O|O-O|0-0-0|0-0|(?:Кр|[KQRBNФЛСК])[a-h]?[1-8]?x?[a-h][1-8]\"\n  + \"(?:=[QRBNФЛСК])?|[a-h]x[a-h][1-8](?:=[QRBNФЛСК])?)[+#]?(?![A-Za-z0-9А-Яа-я])\");\n```\n\nYes, the model once wrote an entire invented line in Russian notation, and the English-only regex let it through.\n\nThere are four checks: invented moves, numbered lines that don't match any known line, claims in words like \"the bishop on c4 hits f7\" (checked on the board), and moves named in words in 8 languages. If anything is found, the client gets a `retract` event (the text already streaming on screen is erased), and the model gets:\n\n```\ncheck.append(\"Your answer names moves that are not in the facts or in any analyze result: \")\n     .append(String.join(\", \", invented))\n     .append(\". Check them with the analyze tool or leave them out. Then give the answer again.\");\n```\n\nOne regeneration only, and only clean answers are cached.\n\nThe downside: the check sometimes punishes the truth. After a quiet move, the model wrote \"Ne5 threatens Nxd3\" — a real threat — but Nxd3 wasn't in the facts, so it was flagged and the second answer was worse. The fix was more facts (`threatens_next`, `takes_away`), not a looser check.\n\nClaude 5 thinks before answering, and **thinking counts against `max_tokens`**. I set 1400 (\"plenty for 120 words\") and answers started cutting off mid-word. On a hard question, the model thought for 45 seconds and returned nothing.\n\n```\n// Claude 5 models think before they answer, and the thinking counts against max_tokens: at 1400\n// an answer of some 450 letters was cut off mid-word.\nstatic final int MAX_TOKENS = 4000;\nstatic final String DEFAULT_EFFORT = \"low\";   // output_config.effort\n```\n\n| effort | question 1 | question 2 | question 3 | \n|---|---|---|---|\n| default | 7.5 s | 10.4 s | 53 s, empty answer | \n| `low` | 12 s | 5.5 s | 8 s | \n\nAnswers stream over SSE (`text` / `retract` / `done`); first words arrive in 2.5–4 s. The system prompt (25 numbered rules, each born from a bug) is split into a cached block with `cache_control: ephemeral` and a per-question `<facts>` block.\n\nWhile teaching the coach not to lie, I found the game review it relied on had been lying for a year:\n\n```\n} else if (lastCp != null) {\n    // UCI scores are relative to the side to move\n    out.cp = whiteToMove ? lastCp : -lastCp;\n```\n\nThe old code read UCI scores as White's, so the eval flipped sign every other ply and half of every eval graph on the site was upside down. Mate scores were dropped entirely. Time per move had been saved as zero for every move ever played. Per-game accuracy now uses Lichess's formulas and a harmonic mean (a plain mean gave \"23 accurate moves and 3 blunders = 83%\" for a lost game).\n\nThe most useful feature turned out to be the least flashy. Over the last 30 games, code looks for circumstances where a habitual mistake happens more often than your baseline:\n\n```\ndouble share = (double) inHabit / habitTotal;   // share among the mistakes\ndouble base  = (double) inAll / allTotal;       // share among all moves\ndouble lift  = share / base;\nif (share < 0.15 || lift < 1.3) return;\n```\n\n\"Most of your blunders are in the opening\" means nothing if most of your moves are in the opening. Only facts that pass both thresholds reach the model, which phrases them. My own result: 66% of my blunders happen on a move where I captured something, against a 29% baseline. Capture, feel good, stop looking.\n\nElevenLabs. Only v4, v4 Turbo, and v3 cover all 22 of the site's languages (Multilingual v2 lacks Bengali, Persian, Hebrew, Vietnamese). Live talk uses `eleven_v4_turbo`; the 33-lesson course is voiced ahead of time with `eleven_v4`, each text paid for once. Everything revolves around caching: short stock phrases live forever in R2 under `sha256(voice|model|text)`, and course audio under `sha256(voice|model|lang|text)` with `immutable`. A semaphore limits concurrent requests to 5 (the plan's limit) — without it, the coach went silent mid-sentence at peak.\n\nNotation has to become words: TTS reads \"Nf3\" as letters; a coach says \"knight f3\". In Russian: `\"18. Qb2 Bxb5\"` → `\"ферзь бэ 2, слон берёт бэ 5\"`, with commas between moves so the voice pauses.\n\nAnd then the fun part. The very first thing the Russian voice did was pronounce *ферзь* (queen) with the wrong vowel and stress — confidently, in a pleasant baritone. A coach that mispronounces the queen inspires about as much trust as a surgeon who can't pronounce \"scalpel\".\n\nI can check Russian, Ukrainian, and English by ear. The coach also speaks Hindi, Korean, Persian and Tagalog, and I'm fairly sure a native Bengali speaker is listening to my coach explain a knight fork right now and quietly laughing. If you speak one of those languages, tell me in the comments how it sounds.\n\nSpeech-to-text is free: the Web Speech API in browsers (absent in Firefox, so the mic button hides there) and the system recognizer in the Android app, because Android WebView has no Web Speech API at all. Tagalog needs the `fil-PH` locale, not `tl` — that one cost me an evening.\n\nWhen you play the coach, a weaker coach must not be a *careless* one. Stockfish in the browser computes the top 6 moves (MultiPV 6), and the coach picks randomly among moves within a loss window of 120 / 80 / 45 / 25 / 0 centipawns depending on level. A piece is worth 300+, so it's never inside the window.\n\nTo judge your move, it's compared with the best move **inside the same MultiPV search**. Two separate searches at depth 12 differ by half a pawn of noise — enough to call a normal 7.O-O an \"inaccuracy\".\n\n| Item | Cost | \n|---|---|\n| Model | Claude Sonnet 5, $2 / $10 per 1M tokens (in / out) | \n| Average question | ~2.9k tokens in, ~400 out ≈ **$0.01** | \n| 10 minutes of voice conversation | $0.06–0.20 (vs ~$0.80 for an off-the-shelf voice agent) | \n| Time to first words | 2.5–4 s | \n\nThe takeaway: \"the engine calculates, the model explains, the synthesizer speaks\" works. It works not because of the model, but because the code around it doesn't trust it.\n\nYou can try the coach at [democraticchess.com](https://www.democraticchess.com/coach-game) — guests get one short lesson without signing up.", "url": "https://wpnews.pro/news/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks", "canonical_source": "https://dev.to/daniel_34e6cb821/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-elevenlabs-speaks-2a0n", "published_at": "2026-10-08 19:12:41+00:00", "updated_at": "2026-10-08 19:20:03.309349+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "generative-ai"], "entities": ["Stockfish", "Claude", "Anthropic", "ElevenLabs", "Java", "Spring", "Cloudflare R2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks", "markdown": "https://wpnews.pro/news/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks.md", "text": "https://wpnews.pro/news/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks.txt", "jsonld": "https://wpnews.pro/news/how-i-built-a-virtual-chess-coach-stockfish-calculates-claude-explains-speaks.jsonld"}}