How I built a virtual chess coach: Stockfish calculates, Claude explains, ElevenLabs speaks A developer built a virtual chess coach that pairs Stockfish for move calculation with Claude for natural-language explanations and ElevenLabs for spoken output in 22 languages. The system feeds the LLM pre-computed engine facts rather than raw FEN positions and validates every move, line, and evaluation against the board, regenerating answers once when validation fails. The coach runs Stockfish both in-browser via WASM and as a server-side pool of three native processes with a 3,000-position LRU cache. Every time I lose a game and open the engine analysis, I get the same thing: −2.3 and an arrow. Fine, my move was bad. But why ? What should I have seen? Stockfish knows the answer better than any grandmaster — it just can't say it in words. So I built a coach that sees the board like Stockfish, explains like a human, and talks to you on a video-call-style screen. The first version came together quickly. Teaching it not to lie took much longer. This post covers the architecture, how I feed facts to the LLM instead of positions, how every answer gets validated against the board, voice in 22 languages, costs, and the bugs. There were more bugs than code. LLMs are bad at chess. They lose track of the board, move pieces that aren't there, and confidently evaluate positions they never calculated. For a coach, that's worse than silence: a 1300 player won't notice that "the knight on f3 is defended by the bishop" is false, and will remember the wrong idea. So the one rule that never changed: Every move, line, and evaluation in an answer comes from the engine. The model only explains. Browser Server Java / Spring ─────── ────────────────────── Stockfish WASM Web Worker Pool of native Stockfish 3 processes ├─ coach's own moves MultiPV 6 ├─ game review depth 14, recheck at 18 └─ judging the player's move depth 12 └─ answering questions depth 14, top 3 moves │ facts JSON Question text / voice ─────────────► Claude Sonnet 5 + "analyze" tool │ text Answer validation ──► one regeneration │ Avatar + audio ◄──────────── mp3 ◄──── ElevenLabs v4 Turbo + R2 cache Stockfish runs in two places. In the browser WASM in a Web Worker for anything that must feel instant and cost nothing: the coach's moves when you play against it, and judging your move during the game. On the server as a pool of native processes. The first version spawned one Stockfish per question — half a second to start, which seemed fine until ten questions arrived at once; the small instance ran out of memory, and the whole site went down. Now it's a pool of 3 64 MB hash each , a 20-second queue timeout, and an LRU cache of 3,000 positions, because the same opening positions get asked about constantly. My first version sent the FEN and asked: "explain why this move is bad". The model stared at r1bqkb1r/pppp1ppp/2n2n2/... and made things up. Once it said 6...Nd4 "attacks the knight on f3 and the bishop on e2". There was a queen on e2. Now the model gets pre-computed facts: { "current": { "move": "14...Bxe4", "eval before": "about equal", "eval after": "White is winning about +2.5, a piece up ", "engine best": "14...Re8", "engine reason": "after Bxe4 the reply Nxe4 wins the bishop: it has no retreat", "reply move": "15.Nxe4", "engine top before": "Re8", "h6", "Qd7" }, "habit tags": "hanging piece" } Two non-obvious lessons: +0.2 the model sometimes flipped it to the other side. My favorite line from the logs: When the player asks "what if I take with the rook?", the model has a tool: { "name": "analyze", "input schema": { "properties": { "start": { "enum": "before", "after" }, "moves": { "type": "array", "items": { "type": "string" } }, "label": { "type": "string" } } } } The server plays the moves on a real board max 12 plies , runs Stockfish, and returns the eval, a 6-ply best line, and what the last move attacks. An illegal move returns an error with a note: "Do not guess why; say only that the move is not possible there." That note exists because in live games analyze once defaulted to the position before the player's move, so a perfectly legal move came back "illegal" — and the coach confidently explained: "c3 is taken by your knight." Even with facts and a tool, the model sometimes named moves that existed nowhere. So every answer is checked before the player sees it. All moves that appeared anywhere the game, engine lines, analyze results go into an allowed set, and every move is pulled out of the answer: // Russian piece letters too Кр, Ф, Л, С, К : the model slipped a whole invented line past the // check once by writing "Фxg5". private static final Pattern MOVE = Pattern.compile " ?< A-Za-z0-9А-Яа-я O-O-O|O-O|0-0-0|0-0| ?:Кр| KQRBNФЛСК a-h ? 1-8 ?x? a-h 1-8 " + " ?:= QRBNФЛСК ?| a-h x a-h 1-8 ?:= QRBNФЛСК ? + ? ? A-Za-z0-9А-Яа-я " ; Yes, the model once wrote an entire invented line in Russian notation, and the English-only regex let it through. There are four checks: invented moves, numbered lines that don't match any known line, claims in words like "the bishop on c4 hits f7" checked on the board , and moves named in words in 8 languages. If anything is found, the client gets a retract event the text already streaming on screen is erased , and the model gets: check.append "Your answer names moves that are not in the facts or in any analyze result: " .append String.join ", ", invented .append ". Check them with the analyze tool or leave them out. Then give the answer again." ; One regeneration only, and only clean answers are cached. The downside: the check sometimes punishes the truth. After a quiet move, the model wrote "Ne5 threatens Nxd3" — a real threat — but Nxd3 wasn't in the facts, so it was flagged and the second answer was worse. The fix was more facts threatens next , takes away , not a looser check. Claude 5 thinks before answering, and thinking counts against max tokens . I set 1400 "plenty for 120 words" and answers started cutting off mid-word. On a hard question, the model thought for 45 seconds and returned nothing. // Claude 5 models think before they answer, and the thinking counts against max tokens: at 1400 // an answer of some 450 letters was cut off mid-word. static final int MAX TOKENS = 4000; static final String DEFAULT EFFORT = "low"; // output config.effort | effort | question 1 | question 2 | question 3 | |---|---|---|---| | default | 7.5 s | 10.4 s | 53 s, empty answer | | low | 12 s | 5.5 s | 8 s | Answers stream over SSE text / retract / done ; first words arrive in 2.5–4 s. The system prompt 25 numbered rules, each born from a bug is split into a cached block with cache control: ephemeral and a per-question