{"slug": "training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant", "title": "Training a wake word that actually fires: lessons from an offline voice assistant", "summary": "A developer built an offline wake-word detector for a Raspberry Pi 4 voice assistant named \"Nova\" using openWakeWord and Piper, finding that training positives with Spanish TTS voices instead of English was the single biggest recall improvement. The engineer also traced a total detection failure to a window-size mismatch: openWakeWord expects 1280-sample (80 ms) windows at 16 kHz, while the capture pipeline was feeding 480-sample (30 ms) frames, yielding scores near zero until frames were buffered.", "body_md": "I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.\n\nThe very first thing it has to do is also the easiest to underestimate: notice when you say its\n\nname. If the **wake word** doesn't fire, nothing else in the pipeline ever gets a turn — the\n\nspeech-to-text, the LLM, the voice, all of it sits there waiting.\n\nMy wake word is \"Nova\". Getting it to trigger reliably took me down a rabbit hole, and most of what\n\nactually moved the needle was *not* where I expected. This post is the recipe I wish I'd had — with\n\nthe specific mistakes that cost me recall, so you can skip them.\n\nEverything here uses open-source tools: [openWakeWord](https://github.com/dscripka/openWakeWord)\n\nfor the model and [Piper](https://github.com/rhasspy/piper) for synthetic speech. The approach\n\nworks for any keyword, not just mine.\n\nopenWakeWord doesn't train a giant speech model. It trains a **small classifier** on top of a\n\nshared, pre-trained audio embedding. The pipeline is:\n\n```\naudio → melspectrogram → shared embedding → small \"is this the wake word?\" classifier\n```\n\nThat's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the\n\ntiny classifier is yours. You feed it two things: **positives** (lots of people saying your word)\n\nand **negatives** (speech and noise that is *not* your word). The catch is that you rarely have\n\nthousands of real recordings of a made-up name — so you synthesize them.\n\nThe standard trick is to generate thousands of utterances of your wake word with a text-to-speech\n\nengine, varying voice, speed and pitch. I used Piper.\n\nHere is the lesson, and it's the big one: **the language of the TTS voice has to match how you'll actually say the word.** My first model was trained with English voices. In English, \"Nova\" is\n\nSwitching the positives to **Spanish Piper voices** was the single biggest improvement I made.\n\nAfter listening to a bunch of candidates, I kept the three that pronounced \"Nova\" cleanly\n\n(`es_ES-sharvard-medium`, `es_MX-ald-medium`, `es_MX-claude-high`) and dropped the ones that sounded\n\noff on this particular word.\n\n```\n# Grab a few Spanish voices and generate a small batch to listen to FIRST\npython -m piper.download_voices --download-dir voices \\\n    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high\n# then synthesize many \"nova\" clips at 16 kHz mono, varying voice/speed/pitch\n```\n\n**Do this before you train anything:** generate ~50 clips and actually *listen*. One minute of\n\nlistening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you\n\nhear in those clips is what the model is about to learn.\n\nPositives alone teach the model to say \"yes\"; it also has to learn to say \"no\". openWakeWord's\n\ntraining uses large public datasets of general speech and noise as negatives, plus pre-computed\n\nfeatures to validate false positives. On top of that it **augments** the positives by mixing in\n\nnoise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.\n\nYou mostly get this for free from the project's training notebook — just don't skip it.\n\nTraining downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google\n\nColab and Kaggle both work and are free. The config change that makes it *your* word is a single\n\nline:\n\n```\ntarget_phrase = [\"nova\"]\n```\n\nTwo hardware gotchas that wasted my time:\n\nOut comes a single `nova.onnx` file. That's the whole model.\n\nWith the model in place, it still wouldn't fire — and this one had nothing to do with training.\n\nopenWakeWord expects to be fed audio in **1280-sample windows (80 ms at 16 kHz)**. My audio capture\n\nwas handing it **480-sample frames (30 ms)**, because that's the frame size the voice-activity\n\ndetector wanted. Fed the wrong window size, the detector returned scores near **0** on every frame\n\nand never triggered. The fix was to buffer incoming frames up to 1280 samples before calling\n\n`predict()`:\n\n```\n# accumulate frames until we have a full 1280-sample window, then score\nbuffer.extend(frame)\nwhile len(buffer) >= 1280:\n    window, buffer = buffer[:1280], buffer[1280:]\n    score = model.predict(window)\n```\n\nIf your freshly trained model scores zero on *everything*, suspect the plumbing before you blame the\n\ntraining.\n\nEarly on I \"tested\" the wake word by saying it a few times and nodding. That's how you fool\n\nyourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of\n\nnegatives, and a sweep across thresholds that prints recall and false positives per threshold.\n\n```\nevaluate_wakeword --positives eval/positives --negatives eval/negatives \\\n    --thresholds 0.2,0.3,0.35,0.4,0.5\n```\n\nUse **at least ~30 varied clips** — different distances, speeds, background noise. In my set, about\n\na dozen of them were specifically the kind that fool a naive model, and they're the ones that tell\n\nyou the truth.\n\nThis also killed a tempting assumption: that the **detection threshold** is the lever for recall.\n\nIt isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no\n\nthreshold saves you — in fact lowering it from 0.4 to 0.3 made things *worse* (more false\n\npositives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.\n\nThe synthetic-only model learned a TTS \"Nova\", not *my* voice in *my* room through *my* microphone.\n\nThe biggest, most reliable win was adding **real recordings** of me saying the word and retraining.\n\nI wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,\n\nspeed and background noise across ~60 of them (and held ~10% back for evaluation). In the training\n\nnotebook, those real positives get **upsampled** and mixed in with the synthetic ones.\n\nThe result, measured on the same evaluation set:\n\n**Recall went from 53% → 70% at threshold 0.35**, just by folding in real-voice positives.\n\nNot magic, but a real, measured step up — and exactly the kind of improvement that's invisible if\n\nyou're only judging by ear.\n\nHere's the minimal path with public tools, so you can train your own keyword. Swap `\"nova\"` for\n\nyours throughout.\n\n**1. Install and synthesize positives.** Generate a few thousand clips, varying voice and speed:\n\n```\npip install piper-tts openwakeword\npython -m piper.download_voices --download-dir voices \\\n    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high\npython\nimport random, subprocess, pathlib\npathlib.Path(\"positives\").mkdir(exist_ok=True)\nvoices = [\"es_ES-sharvard-medium\", \"es_MX-ald-medium\", \"es_MX-claude-high\"]\nfor i in range(2000):\n    v = random.choice(voices)\n    length = round(random.uniform(0.9, 1.3), 2)   # speed variation\n    subprocess.run(\n        [\"piper\", \"--model\", f\"voices/{v}.onnx\",\n         \"--length-scale\", str(length), \"--output_file\", f\"positives/{v}_{i}.wav\"],\n        input=b\"nova\\n\",\n    )\n# (CLI flags vary slightly by Piper version — check `piper --help`.)\n```\n\nThen **listen to a handful** before going further.\n\n**2. Train.** Open openWakeWord's `automatic_model_training.ipynb`\n\n([in their repo](https://github.com/dscripka/openWakeWord)) on Kaggle or Colab, select a **T4**\n\nGPU, set `target_phrase = [\"nova\"]`, point it at your `positives/` folder, and *Run All*. It pulls\n\nthe negative/background datasets and augmentation for you and exports a single `nova.onnx`.\n\n**3. Evaluate** against your own clip folders and sweep thresholds — this is the step people skip:\n\n``` python\nimport os, numpy as np, soundfile as sf\nfrom openwakeword.model import Model\n\nmodel = Model(wakeword_models=[\"nova.onnx\"])\n\ndef best_score(path):\n    model.reset()\n    audio, _ = sf.read(path)                       # 16 kHz mono\n    audio = (audio * 32767).astype(np.int16)\n    top = 0.0\n    for i in range(0, len(audio) - 1280, 1280):    # 80 ms windows\n        top = max(top, model.predict(audio[i:i + 1280])[\"nova\"])\n    return top\n\npos = [best_score(f\"eval/positives/{f}\") for f in os.listdir(\"eval/positives\")]\nneg = [best_score(f\"eval/negatives/{f}\") for f in os.listdir(\"eval/negatives\")]\nfor t in (0.2, 0.3, 0.35, 0.4, 0.5):\n    recall = sum(s >= t for s in pos) / len(pos)\n    fp = sum(s >= t for s in neg)\n    print(f\"thr={t}: recall={recall:.0%}  false_positives={fp}/{len(neg)}\")\n```\n\n**4. Fine-tune with real voice (optional, high impact).** Record ~60 clips of yourself saying the\n\nword (16 kHz mono — a few lines with `sounddevice`), hold back ~10% for evaluation, and add them to\n\nthe training set with upsampling. This is what took me from 53% to 70%.\n\n**5. Deploy.** Copy `nova.onnx` to your device and feed the detector **1280-sample windows** (see\n\nthe buffering snippet above). Tune the threshold from your evaluation numbers.\n\nIf you only remember four things:\n\nAnd one bonus, because it bit me hardest: if a model scores zero on everything, check the window\n\nsize you're feeding it before you retrain anything.\n\nHave you trained a custom wake word? I'd love to hear what moved recall for you — especially for\n\nnon-English keywords.\n\n**Want more context, or to see how we did it?**\n\nNova isn't fully public yet, but you can get **early access** to the repository — all the\n\ndocumentation and the complete source code — at **[https://gitlab.com/gabrielhruiz1/nova](https://gitlab.com/gabrielhruiz1/nova)**.\n\nLeave us a message and we'll try to grant you access ASAP, until we publish everything\n\nofficially (we're still working on a few parts).", "url": "https://wpnews.pro/news/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant", "canonical_source": "https://dev.to/gabrielhruiz/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant-3716", "published_at": "2026-10-08 11:14:26+00:00", "updated_at": "2026-10-08 11:19:18.529807+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "machine-learning", "developer-tools"], "entities": ["openWakeWord", "Piper", "Raspberry Pi 4", "Google Colab", "Kaggle", "Nova"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant", "markdown": "https://wpnews.pro/news/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant.md", "text": "https://wpnews.pro/news/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant.txt", "jsonld": "https://wpnews.pro/news/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant.jsonld"}}