Training a wake word that actually fires: lessons from an offline voice assistant A developer built an offline wake-word detector for a Raspberry Pi 4 voice assistant named "Nova" using openWakeWord and Piper, finding that training positives with Spanish TTS voices instead of English was the single biggest recall improvement. The engineer also traced a total detection failure to a window-size mismatch: openWakeWord expects 1280-sample (80 ms) windows at 16 kHz, while the capture pipeline was feeding 480-sample (30 ms) frames, yielding scores near zero until frames were buffered. I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4. The very first thing it has to do is also the easiest to underestimate: notice when you say its name. If the wake word doesn't fire, nothing else in the pipeline ever gets a turn — the speech-to-text, the LLM, the voice, all of it sits there waiting. My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what actually moved the needle was not where I expected. This post is the recipe I wish I'd had — with the specific mistakes that cost me recall, so you can skip them. Everything here uses open-source tools: openWakeWord https://github.com/dscripka/openWakeWord for the model and Piper https://github.com/rhasspy/piper for synthetic speech. The approach works for any keyword, not just mine. openWakeWord doesn't train a giant speech model. It trains a small classifier on top of a shared, pre-trained audio embedding. The pipeline is: audio → melspectrogram → shared embedding → small "is this the wake word?" classifier That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the tiny classifier is yours. You feed it two things: positives lots of people saying your word and negatives speech and noise that is not your word . The catch is that you rarely have thousands of real recordings of a made-up name — so you synthesize them. The standard trick is to generate thousands of utterances of your wake word with a text-to-speech engine, varying voice, speed and pitch. I used Piper. Here is the lesson, and it's the big one: the language of the TTS voice has to match how you'll actually say the word. My first model was trained with English voices. In English, "Nova" is Switching the positives to Spanish Piper voices was the single biggest improvement I made. After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly es ES-sharvard-medium , es MX-ald-medium , es MX-claude-high and dropped the ones that sounded off on this particular word. Grab a few Spanish voices and generate a small batch to listen to FIRST python -m piper.download voices --download-dir voices \ es ES-sharvard-medium es MX-ald-medium es MX-claude-high then synthesize many "nova" clips at 16 kHz mono, varying voice/speed/pitch Do this before you train anything: generate ~50 clips and actually listen . One minute of listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you hear in those clips is what the model is about to learn. Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's training uses large public datasets of general speech and noise as negatives, plus pre-computed features to validate false positives. On top of that it augments the positives by mixing in noise and room impulse responses RIRs so the model survives a real room instead of only clean TTS. You mostly get this for free from the project's training notebook — just don't skip it. Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google Colab and Kaggle both work and are free. The config change that makes it your word is a single line: target phrase = "nova" Two hardware gotchas that wasted my time: Out comes a single nova.onnx file. That's the whole model. With the model in place, it still wouldn't fire — and this one had nothing to do with training. openWakeWord expects to be fed audio in 1280-sample windows 80 ms at 16 kHz . My audio capture was handing it 480-sample frames 30 ms , because that's the frame size the voice-activity detector wanted. Fed the wrong window size, the detector returned scores near 0 on every frame and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling predict : accumulate frames until we have a full 1280-sample window, then score buffer.extend frame while len buffer = 1280: window, buffer = buffer :1280 , buffer 1280: score = model.predict window If your freshly trained model scores zero on everything , suspect the plumbing before you blame the training. Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of negatives, and a sweep across thresholds that prints recall and false positives per threshold. evaluate wakeword --positives eval/positives --negatives eval/negatives \ --thresholds 0.2,0.3,0.35,0.4,0.5 Use at least ~30 varied clips — different distances, speeds, background noise. In my set, about a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell you the truth. This also killed a tempting assumption: that the detection threshold is the lever for recall. It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no threshold saves you — in fact lowering it from 0.4 to 0.3 made things worse more false positives, no real gain . The threshold is a fine-tuning dial, not a fix for bad training data. The synthetic-only model learned a TTS "Nova", not my voice in my room through my microphone. The biggest, most reliable win was adding real recordings of me saying the word and retraining. I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone, speed and background noise across ~60 of them and held ~10% back for evaluation . In the training notebook, those real positives get upsampled and mixed in with the synthetic ones. The result, measured on the same evaluation set: Recall went from 53% → 70% at threshold 0.35 , just by folding in real-voice positives. Not magic, but a real, measured step up — and exactly the kind of improvement that's invisible if you're only judging by ear. Here's the minimal path with public tools, so you can train your own keyword. Swap "nova" for yours throughout. 1. Install and synthesize positives. Generate a few thousand clips, varying voice and speed: pip install piper-tts openwakeword python -m piper.download voices --download-dir voices \ es ES-sharvard-medium es MX-ald-medium es MX-claude-high python import random, subprocess, pathlib pathlib.Path "positives" .mkdir exist ok=True voices = "es ES-sharvard-medium", "es MX-ald-medium", "es MX-claude-high" for i in range 2000 : v = random.choice voices length = round random.uniform 0.9, 1.3 , 2 speed variation subprocess.run "piper", "--model", f"voices/{v}.onnx", "--length-scale", str length , "--output file", f"positives/{v} {i}.wav" , input=b"nova\n", CLI flags vary slightly by Piper version — check piper --help . Then listen to a handful before going further. 2. Train. Open openWakeWord's automatic model training.ipynb in their repo https://github.com/dscripka/openWakeWord on Kaggle or Colab, select a T4 GPU, set target phrase = "nova" , point it at your positives/ folder, and Run All . It pulls the negative/background datasets and augmentation for you and exports a single nova.onnx . 3. Evaluate against your own clip folders and sweep thresholds — this is the step people skip: python import os, numpy as np, soundfile as sf from openwakeword.model import Model model = Model wakeword models= "nova.onnx" def best score path : model.reset audio, = sf.read path 16 kHz mono audio = audio 32767 .astype np.int16 top = 0.0 for i in range 0, len audio - 1280, 1280 : 80 ms windows top = max top, model.predict audio i:i + 1280 "nova" return top pos = best score f"eval/positives/{f}" for f in os.listdir "eval/positives" neg = best score f"eval/negatives/{f}" for f in os.listdir "eval/negatives" for t in 0.2, 0.3, 0.35, 0.4, 0.5 : recall = sum s = t for s in pos / len pos fp = sum s = t for s in neg print f"thr={t}: recall={recall:.0%} false positives={fp}/{len neg }" 4. Fine-tune with real voice optional, high impact . Record ~60 clips of yourself saying the word 16 kHz mono — a few lines with sounddevice , hold back ~10% for evaluation, and add them to the training set with upsampling. This is what took me from 53% to 70%. 5. Deploy. Copy nova.onnx to your device and feed the detector 1280-sample windows see the buffering snippet above . Tune the threshold from your evaluation numbers. If you only remember four things: And one bonus, because it bit me hardest: if a model scores zero on everything, check the window size you're feeding it before you retrain anything. Have you trained a custom wake word? I'd love to hear what moved recall for you — especially for non-English keywords. Want more context, or to see how we did it? Nova isn't fully public yet, but you can get early access to the repository — all the documentation and the complete source code — at https://gitlab.com/gabrielhruiz1/nova https://gitlab.com/gabrielhruiz1/nova . Leave us a message and we'll try to grant you access ASAP, until we publish everything officially we're still working on a few parts .