Open Codenames: a new way to benchmark LLMs Open Codenames is a new open-source Codenames variant that pits a human player against and alongside language models, released under the MIT license and requiring Node.js 22 or newer. The game runs on Windows desktop via Electron and in the browser, supports bring-your-own API keys from Anthropic, OpenAI, Google, OpenRouter, xAI, Mistral, or a local model, and ships with a built-in deck of 68 public-domain engravings by J. J. Grandville. Keys stay with the player and go only to the chosen provider because there is no Open Codenames server. Codenames with pictures, played with AI teammates and opponents. You and an AI teammate face two AI opponents on a board of pictures. Give one-word clues as the spymaster, or work out your teammate's clues as the operative. The AI players are language models that look at the pictures themselves, reason out loud in the game log, and never repeat a clue. A debrief after every game shows what each clue meant and what the AIs considered. - Desktop Windows and browser versions. - Bring your own key: Anthropic, OpenAI, Google, OpenRouter, xAI, Mistral, or a local model on the desktop. Keys stay with the player and go only to the provider they chose; there is no Open Codenames server. - Built-in picture deck: 68 public-domain engravings by J. J. Grandville. The desktop version can also generate new pictures or use your own collections. Player documentation: docs/README.md https://github.com/attilatorda/open-codenames/blob/main/docs/README.md also shipped with the game . Requires Node.js 22 or newer. npm install npm run dev desktop game with hot reload debug build, offline mock AI available npm run dev:web browser version npm test See CONTRIBUTING.md https://github.com/attilatorda/open-codenames/blob/main/CONTRIBUTING.md to get involved and SECURITY.md https://github.com/attilatorda/open-codenames/blob/main/SECURITY.md to report a vulnerability. MIT https://github.com/attilatorda/open-codenames/blob/main/LICENSE . The picture deck is public domain; fonts and bundled libraries are listed in THIRD PARTY NOTICES.md https://github.com/attilatorda/open-codenames/blob/main/THIRD PARTY NOTICES.md . “Codenames” is a registered trademark of Czech Games Edition. Open Codenames is an independent fan project and is not affiliated with or endorsed by Czech Games Edition. Electron + TypeScript + React, local-first. Player-facing docs live in docs/ https://github.com/attilatorda/open-codenames/blob/main/docs and ship with the build. Two flavors, fixed at build time by OC BUILD src/shared/build.ts : - Debug default for dev / build : includes the offline mock AI and placeholder image generator, shows a DEBUG badge, keeps its data in %APPDATA%\Open Codenames Debug . - Release OC BUILD=release : no mock code at all. This is what ships on itch.io. Two targets: desktop Electron and web the same UI in a browser; src/web replaces the main process with an in-page bridge — settings, keys and history in local storage, the standard deck as static files, provider calls straight from the page; no picture generation, folders, imports or local models . | Command | What it does | |---|---| | npm run dev | Run the desktop game debug with hot reload | | npm run dev:web | Run the web version debug in a browser with hot reload | | npm test | Unit tests Vitest : rules, information boundaries, strategy, parsing, providers, credentials | | npm run typecheck | Type-check main, preload, core, renderer and web | | npm run e2e | Build debug , then drive the desktop app with Playwright against the mock AI; screenshots go to .e2e/ | | npm run e2e:web | Build the debug web version, then drive it in headless Chrome; screenshots go to .e2e-web/ | | npm run dist:debug | Debug desktop build → release/debug/win-unpacked/Open Codenames Debug.exe | | npm run dist:release | Release desktop build → release/desktop/ zip with .itch.toml and docs | | npm run build:web | Release web build → out/web/ | | npm run dist:itch | Both release builds, packaged for itch.io in release/itch/ Windows zip, web zip, UPLOAD.txt | | npm run demos | Record the two ~30 s demo videos 1920×1080 MP4, captioned into release/demos/ ; needs ffmpeg and Chrome | | python -P scripts/deck/fetch met.py , fetch commons.py , survey.py , cut.py , build deck.py | Rebuild the standard deck from public-domain sources selection and captions in scripts/deck/selection.json | | npx tsx scripts/mock-sim.mts games seed | Watch the offline mock AI play itself on standard-deck boards clues, intended pictures, picks | | Play Open Codenames.cmd | Launch the debug desktop build | If your shell sets ELECTRON RUN AS NODE=1 VS Code’s terminal does , unset it before launching Electron directly. src/ core/ Pure TypeScript game logic — no Electron, React or vendor code. engine/ GameEngine authoritative rules/state , views information boundaries , MatchController turn loop rules/ IRuleSet + StandardRules board/ Layouts, concept bank, board generator seeded ai/ LLMPlayer, the AI profile, prompts, strategy clue size, risk/reward, stopping , JSON parsing, mock brain players/ IPlayer, HumanPlayer UI bridge modes/ Game-mode registry Quick Play, Standard, AI vs AI, coming-soon modes replay/ GameRecord + recorder debriefs, future replays / AI Arena images/ Prompt construction and art styles shared/ Types shared by main and renderer: settings, provider catalog, IPC contract main/ Electron main process: the only place keys and network calls exist providers/llm/ ILLMProvider: anthropic official SDK , openai Responses API , google Interactions API , openaiCompatible OpenRouter, xAI, Mistral, local , mock providers/image/ IImageGenerationProvider: stability, bfl, openaiImages, googleImages, replicate, a1111, mock images/ Image library disk cache + player folder , thumbnails for vision, board image pipeline store/ Settings, CredentialStore safeStorage/DPAPI , replays, diagnostics redacting log preload/ Typed contextBridge API window.oc — keys can be written but never read back renderer/ React UI: screens, board, player boards, game log, debrief web/ Browser version of the window.oc bridge entry for the web build Dependency direction: renderer and main depend on core / shared ; core depends on nothing app-specific. buildSpymasterView and buildOperativeView src/core/engine/views.ts are the only way players see the game. OperativeView has no field that can carry the key, prompt builders accept only the view for their role, and tests/boundaries.test.ts checks non-interference: games that differ only in hidden identities produce byte-identical operative prompts. The UI renders the board from the human’s role view too. Every AI seat is its own LLMPlayer with its own slot, RNG and memory, all playing the single profile in src/core/ai/personalities.ts . Players see the board pictures themselves thumbnails ; no view or prompt carries a text description of a picture, so only vision models can take a seat. Spymaster: clue size — the human teammate picks it in the status bar before each clue Auto or 1–4 , otherwise planClueSize opening 1/2/3 at 25/50/25%, then “stay ahead” of the opponents’ pace → LLM proposes candidate clues with link strengths and risks → evaluateCandidates validates them against the real key, rejects any clue already given this game, and scores expected value → teammate simulation → clue. Operative: one LLM call per turn ranks guesses with confidences → shouldTakeGuess decides how far to go given the score. - LLM provider : implement ILLMProvider in src/main/providers/llm/ , register it in registry.ts , add catalog data in src/shared/providers.ts . - Image provider : implement IImageGenerationProvider in src/main/providers/image/ , register it, add catalog data. - Game mode : add a GameModeDef in src/core/modes/gameModes.ts . - Rule variant : implement IRuleSet and pass it to GameEngine . %APPDATA%\Open Codenames Electron userData : settings.json , credentials.json encrypted , images/ , replays/ , logs/ . Set OC USER DATA to use a different folder the e2e test does .