{"slug": "voice-racer-game-built-with-oss-model-neudecide", "title": "Voice racer game built with OSS model NeuDecide", "summary": "Neuphonic's 43 MB NeuDecide speech-to-tool-call model powers Voice Racer, a pixel-art racing game driven by spoken commands without any speech-to-text or chatbot in the loop, according to the project's documentation. The game runs the model locally via a JavaScript port of the neudecide Python package on ONNX Runtime — onnxruntime-node natively or onnxruntime-web (WASM, in a Web Worker) in the browser — and requires Node 22 or newer, a microphone, and desktop Chrome or Edge. A grammar-constrained decoder emits only valid tool-call JSON from a menu of 9 game commands plus noAction, and a confidence gate drops unsure calls so most chatter does nothing.", "body_md": "A pixel-art racing game you drive with your voice. Say **\"go left\"**, **\"speed up\"** or **\"jump\"** and the car does it.\n\nThere is no speech-to-text and no chatbot in the loop. [NeuDecide](https://huggingface.co/neuphonic/neudecide), a small\n(43 MB) speech-to-tool-call model by Neuphonic, hears your voice and picks a game command directly. It runs on your own\nmachine: natively next to the page, or entirely inside the browser with WebAssembly.\n\n<sub>Six spoken commands, decided by the native engine. The voice is a macOS text-to-speech voice played into the\nbrowser as its microphone. A GIF has no sound, so the caption shows what was said.</sub>\n\n``` php\nflowchart LR\n  mic[\"Microphone\"] --> vad[\"Voice detector<br/>(energy VAD)\"]\n  subgraph nd[\"NeuDecide, run by our JavaScript engine\"]\n    ae[\"Audio encoder\"] --> te[\"Tool encoder\"]\n    menu[\"Tool menu<br/>9 game commands + noAction, as JSON\"] --> te\n    te --> dec[\"Constrained decoder<br/>(only valid tool-call JSON)\"]\n  end\n  vad -- \"16 kHz utterance\" --> ae\n  dec -- \"moveLeft()\" --> gate[\"Confidence gate\"]\n  gate --> game[\"Game\"]\n```\n\n1. The **voice detector** listens to the microphone and cuts out each utterance once you have been silent for 120 ms.\n2. **NeuDecide** reads the audio together with a**menu of tools** (`moveLeft` ,`moveRight` ,`moveToLane` ,`speedUp` ,`slowDown` ,`jump` ,`dodge` ,`start` ,`pause` ,`noAction` , see[`src/voice/tools.js`](https://github.com/vimalananddev/voice-racer/blob/main/src/voice/tools.js) ).\nA grammar-constrained decoder can only write a valid tool call, and the game acts the moment the call is\ndetermined, before decoding even finishes.\n3. A **confidence gate** drops unsure calls and \"no action\" answers, so most chatter does nothing.\n4. The **game** turns the call into an action: change lane, boost, brake, jump, or dodge (the game picks the safest lane).\n\nThe engine that runs the model is a JavaScript port of the `neudecide` Python package (tokenizer, constrained decoding,\ninference loop) on [ONNX Runtime](https://onnxruntime.ai): `onnxruntime-node` for the native engine, `onnxruntime-web` (WASM,\nin a Web Worker) in the browser. More detail in [docs/ARCHITECTURE.md](https://github.com/vimalananddev/voice-racer/blob/main/docs/ARCHITECTURE.md).\n\nYou need **Node 22** or newer, a microphone, and desktop Chrome (or Edge).\n\n```\ngit clone https://github.com/vimalananddev/voice-racer.git\ncd voice-racer\nnpm install\n```\n\n**Download the model** (once, about 43 MB). It is free, but gated on Hugging Face: sign in, accept the terms at\n[https://huggingface.co/neuphonic/neudecide](https://huggingface.co/neuphonic/neudecide), then use the Hugging Face CLI, `hf`:\n\n```\nbrew install hf                  # or: pipx install huggingface_hub   or: uv tool install huggingface_hub\nhf auth login                    # paste a read token from https://huggingface.co/settings/tokens\nhf download neuphonic/neudecide --local-dir models/neudecide\n```\n\nNo CLI? Download `config.json`, `tokenizer.json`, `audio_encoder_q4.onnx`, `tool_encoder_q4.onnx` and\n`decoder_step_q4.onnx` from the model page's **Files and versions** tab into `models/neudecide/`.\n([models/README.md](https://github.com/vimalananddev/voice-racer/blob/main/models/README.md) has the details.) Then:\n\n```\nnpm start\n```\n\n1. Open **[http://localhost:5180](http://localhost:5180)** .\n2. Wait for **Ready!** , click**Turn on the mic** and allow the microphone.\n3. Say **\"start the game\"** , then drive: \"go left\", \"go right\", \"speed up\", \"jump\"...\n\nPressing **Enter** instead of clicking turns on the mic and starts the race in one go (then skip step 3).\n\n`npm start` runs the game page on port 5180 and the native voice engine on `ws://127.0.0.1:5190`, both on this\nmachine only. Ctrl-C stops both. Busy ports? `PORT=5280 VOICE_PORT=5290 npm start` (the printed address then\ncarries `?localport=5290`). `npm run check` runs a quick smoke test that needs no microphone.\n\n`npm run dev` runs only the game page, so the in-browser engine is used. Without a native engine, the browser console\nshows one failed WebSocket connection to `ws://127.0.0.1:5190` per page load: that is the page looking for the native\nengine, and it is expected.\n\n| say | does | hit rate* | \n|---|---|---|\n| \"start the game\" | start, restart after a crash, or resume (Enter works too) | 19/24 | \n| \"go left\" · \"go right\" | one lane left / right | 30/32 · 31/32 | \n| \"lane number three\" | go straight to lane 1-4 | 53/56 | \n| \"speed up\" | boost for ~2.5 s | 24/24 | \n| \"slow down\" | brake for ~2 s | 24/24 | \n| \"jump\" | hop over cones, oil and potholes (not over cars) | 23/24 | \n| \"dodge it\" | the game moves you to the safest lane | 23/24 | \n| \"pause the game\" · \"resume\" | pause / carry on | 21/24 · 24/24 | \n\n* Synthetic voices (macOS text-to-speech), quiet room, no game sound in the microphone, native engine. Real voices and rooms will differ.\n\nTips:\n\n- **Wear headphones** (or press**M** to mute the game). Game sound in the microphone is the main reason\n\"left\" and \"right\" get missed: with it, the one-lane commands drop from about 9 in 10 to about half.\n- Say **\"go left\" / \"go right\"** rather than the bare word. Every left / right moves**one** lane; say it twice to reach\nthe edge lane, or use the lane number (\"lane number four\"). \"Move left\", \"shift left\" and \"switch right\" work too.\n- **Don't say \"alright\" while driving** : the model hears \"right\" and changes lane. \"All right\", \"that's right\" and\n\"yeah right\" do it too. Press**V** to mute the microphone before you chat.\n- When a command is missed, usually nothing happens. If the game pauses by surprise, say \"resume\".\n- Ordinal lanes (\"second lane\", \"last lane\") and slang (\"floor it\") almost never work.\n\nAll latency and hit-rate figures in this README were measured on an M1 Pro MacBook Pro (the browser engine in desktop Chrome), with synthetic voices.\n\n| engine | how to pick it | where it runs | from the end of your speech to the game reacting | \n|---|---|---|---|\n| **Native** (default) | `npm start` , or`?engine=local` | onnxruntime-node next to the page (the tool encoder on 5 threads on an M1 Pro; `min(6, cores - 2)` on other machines) | **~200 ms** (model time ~80-100 ms) | \n| **Browser** | `?engine=browser` , or automatic when no native engine answers | onnxruntime-web (WASM) in a Web Worker | **~0.3-0.6 s** (4 threads: ~0.35 s; ~0.7 s where only 1 thread works) | \n| **Remote** | `?engine=ws://host:5190/?token=…` | the native engine on another machine | that machine's model time + the network | \n\nThe ~200 ms includes the 120 ms the voice detector waits to be sure you have stopped talking (with `?ptt=1` it starts\nwhen you release the talk key instead). Almost all of the model time is the **tool encoder**: it mixes the tool menu\nwith your audio from its first layer, so it cannot be cached and runs for every command. WebGPU is not used: it fails on\nApple GPUs (the model needs 11 storage buffers, the limit is 10).\n\nIf the engine stops mid-game, the page says \"Voice is reconnecting…\" and reconnects by itself. In the default mode, a native engine that is not back within 4 s is replaced by the browser engine (reload to go back). The keyboard always works.\n\nThe engine is `scripts/remote-voice-server.mjs`. It was tested on macOS and on Linux (arm64), and should run anywhere\nonnxruntime-node does. A machine with more CPU cores is faster per command. On that machine, with the repo,\n`npm install` and the model in place:\n\n```\nTOKEN=pick-a-secret HOST=0.0.0.0 PORT=5190 npm run engine\n```\n\nThen open `http://localhost:5180/?engine=ws://<that-machine>:5190/?token=pick-a-secret` on the machine with the game.\nThe `?dev=1` panel names the engine by the hostname it reports; start it with `ENGINE_NAME=desktop` (any name) to show\nsomething else, for example before you record the panel (`?record=1` just says \"Remote\"). Always set a `TOKEN` when\n`HOST` is not `127.0.0.1`, and use it on a network you trust: the link is plain `ws://`. Even simpler and private: keep the engine on `127.0.0.1` there and use an\nSSH tunnel, `ssh -N -L 5195:127.0.0.1:5190 you@that-machine`, then open `?engine=ws://127.0.0.1:5195`.\n\nThe engine prints its thread settings at start-up; `TOOL_THREADS=n` tries others. See the header of\n[`scripts/remote-voice-server.mjs`](https://github.com/vimalananddev/voice-racer/blob/main/scripts/remote-voice-server.mjs) for all options and the protocol.\n\n| option | what it does | \n|---|---|\n| `?dev=1` | the developer panel: the tool call, latency, engine, confidence and a decision log, instead of the clean panel | \n| `?record=1` | a clean layout for screen recording (9:16 window: one column; 4:5, 1:1 or 16:9: two columns), quieter game sound, a faster day-to-night cycle and busier traffic | \n| `?ptt=1` | push-to-talk: hold **T** (or hold the mic button) while you speak | \n| `?engine=local` ·`?engine=browser` ·`?engine=ws://host:port/?token=…` | pick the engine (default: native if it answers, else browser; `native` and`wasm` work too) | \n| `?voice=mock` | no model at all: the keyboard fakes voice events (for working on the UI) | \n| `?minconf=0.6` | a stricter confidence gate (default 0.45; lower lets noise through) | \n| `?tools=fast` | a shorter tool menu: faster, but less accurate | \n| `?tod=fast` ·`?traffic=busy` | the recording pacing outside `?record=1` (`=normal` turns it off inside) | \n| `?tod=night` | start at that time of day ( `day` ,`sunset` ,`dusk` ,`night` or`dawn` ) | \n| `?sfx=0` | no game sound at all | \n| `?ec=0` ·`?ns=0` ·`?agc=0` | turn off the browser's echo cancellation / noise suppression / auto gain on the mic (on by default: they help keep game sound out of the mic) | \n| `?resampler=own` ·`?resampler=browser` | force how the mic is resampled to 16 kHz (default: the browser's own in Chrome / Edge, ours elsewhere) | \n| `?brand=@yourhandle` | a small watermark in the corner | \n\nDeveloper options: `?god=1` (the car cannot crash) and `?keepalive=1` (keeps running in a background tab). The browser\nconsole has more: see [docs/ARCHITECTURE.md](https://github.com/vimalananddev/voice-racer/blob/main/docs/ARCHITECTURE.md#debugging-from-the-console).\n\n<sub>Left: `?dev=1` (real decisions and latencies from a test run). Right: `?record=1` in a 9:16 window.</sub>\n\n| key | does | \n|---|---|\n| **V** | microphone on / off | \n| **T** (with`?ptt=1` ) | hold to talk | \n| ← → or A D | one lane left / right | \n| Shift + ← → | far left / far right (keyboard and touch buttons only) | \n| 1 2 3 4 | go to that lane | \n| ↑ W / ↓ S | boost / brake | \n| Space / X | jump / dodge | \n| Enter / P | start or resume / pause | \n| M | game sound on / off | \n\nOn a touch screen, touch buttons appear under the game. Playing from a phone needs your own HTTPS setup in front of\nthe game server: `server.mjs` listens on 127.0.0.1 only, and browsers allow the microphone only on `https://` or\n`localhost`.\n\n- **English only.**\n- **Tested with synthetic voices.** The numbers above come from macOS text-to-speech voices, not from real people.\nTry it with your own voice and microphone before you rely on it.\n- **Game sound in the microphone hurts a lot** , especially bare \"left\" / \"right\". Use headphones.\n- **Chatter can trigger commands.** Even with the confidence gate, about 1 in 8 chatter or noise clips fired something\nin tests (usually a pause, a dodge or a lane change), and some phrases, like \"thanks for watching\", often fire a dodge.\nMute the microphone (V) or use push-to-talk when you talk to people.\n- The browser engine is about 2x slower than the native one in desktop Chrome, and much slower in browsers that cannot run it multi-threaded. Developed and tested mostly in desktop Chrome on macOS.\n\n```\nindex.html, styles.css        the page\nserver.mjs                    static server (127.0.0.1 only; COOP/COEP headers for multi-threaded WASM)\nscripts/\n  start.mjs                   npm start: the game server + the native engine\n  remote-voice-server.mjs     the native voice engine (WebSocket), also for another machine\n  vendor-ort.mjs              copies onnxruntime-web into vendor/ort/ (runs after npm install)\n  check.mjs                   npm run check: smoke test, no microphone or audio files\nsrc/\n  main.js                     wiring: engine choice, microphone, voice events -> game actions\n  game/                       the racer: world, traffic, sprites and font (all drawn in code), lighting, HUD,\n                              synthesised sound effects\n  ui/                         the voice panel (show-overlay.js; overlay.js is the ?dev=1 panel), layout,\n                              title card, touch buttons, mock voice\n  voice/                      NeuDecide in JavaScript: tokenizer, constrained decoding, inference core, Web Worker,\n                              microphone capture, VAD, confidence gate, tool menu\nmodels/README.md              how to get the model (the weights go in models/neudecide/, never in git)\ndocs/                         ARCHITECTURE.md, screenshots\n```\n\nNo build step and no bundler: plain ES modules served as static files.\n\n- **[NeuDecide](https://huggingface.co/neuphonic/neudecide)** by[Neuphonic](https://www.neuphonic.com) (Apache-2.0): the\nmodel, and the`neudecide` Python package this JavaScript engine is ported from.\n- **[ONNX Runtime](https://onnxruntime.ai)** by Microsoft (MIT):`onnxruntime-web` and`onnxruntime-node` .\n- **[SentencePiece](https://github.com/google/sentencepiece)** by Google (Apache-2.0): the JavaScript tokenizer is ported from it.\n- **[llguidance](https://github.com/microsoft/llguidance)** by Microsoft (MIT): the forced-token logic of the constrained\ndecoder is adapted from its`toktrie` .\n- **[ws](https://github.com/websockets/ws)** (MIT).\n- All game art is original pixel art drawn in code; all sounds are synthesised with Web Audio.\n\nSee [NOTICE](https://github.com/vimalananddev/voice-racer/blob/main/NOTICE) for details.\n\nVoice Racer is an independent project and is not affiliated with or endorsed by Neuphonic. NeuDecide is a model by Neuphonic; the name is used only to identify it.\n\n[Apache License 2.0](https://github.com/vimalananddev/voice-racer/blob/main/LICENSE). Copyright 2026 Vimal Anand.\n\nThe NeuDecide model weights are not part of this repository and are covered by their own terms on Hugging Face.", "url": "https://wpnews.pro/news/voice-racer-game-built-with-oss-model-neudecide", "canonical_source": "https://github.com/vimalananddev/voice-racer", "published_at": "2026-10-11 10:39:10+00:00", "updated_at": "2026-10-11 10:52:51.736805+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-tools", "developer-tools", "natural-language-processing"], "entities": ["Neuphonic", "NeuDecide", "Voice Racer", "ONNX Runtime", "onnxruntime-node", "onnxruntime-web", "Hugging Face", "Node 22"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/voice-racer-game-built-with-oss-model-neudecide", "markdown": "https://wpnews.pro/news/voice-racer-game-built-with-oss-model-neudecide.md", "text": "https://wpnews.pro/news/voice-racer-game-built-with-oss-model-neudecide.txt", "jsonld": "https://wpnews.pro/news/voice-racer-game-built-with-oss-model-neudecide.jsonld"}}