I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.
The very first thing it has to do is also the easiest to underestimate: notice when you say its
name. If the wake word doesn't fire, nothing else in the pipeline ever gets a turn β the
speech-to-text, the LLM, the voice, all of it sits there waiting.
My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what
actually moved the needle was not where I expected. This post is the recipe I wish I'd had β with
the specific mistakes that cost me recall, so you can skip them.
Everything here uses open-source tools: openWakeWord
for the model and Piper for synthetic speech. The approach
works for any keyword, not just mine.
openWakeWord doesn't train a giant speech model. It trains a small classifier on top of a
shared, pre-trained audio embedding. The pipeline is:
audio β melspectrogram β shared embedding β small "is this the wake word?" classifier
That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the
tiny classifier is yours. You feed it two things: positives (lots of people saying your word)
and negatives (speech and noise that is not your word). The catch is that you rarely have
thousands of real recordings of a made-up name β so you synthesize them.
The standard trick is to generate thousands of utterances of your wake word with a text-to-speech
engine, varying voice, speed and pitch. I used Piper.
Here is the lesson, and it's the big one: the language of the TTS voice has to match how you'll actually say the word. My first model was trained with English voices. In English, "Nova" is
Switching the positives to Spanish Piper voices was the single biggest improvement I made.
After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly
(es_ES-sharvard-medium, es_MX-ald-medium, es_MX-claude-high) and dropped the ones that sounded
off on this particular word.
python -m piper.download_voices --download-dir voices \
es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
Do this before you train anything: generate ~50 clips and actually listen. One minute of
listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you
hear in those clips is what the model is about to learn.
Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's
training uses large public datasets of general speech and noise as negatives, plus pre-computed
features to validate false positives. On top of that it augments the positives by mixing in
noise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.
You mostly get this for free from the project's training notebook β just don't skip it.
Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google
Colab and Kaggle both work and are free. The config change that makes it your word is a single
line:
target_phrase = ["nova"]
Two hardware gotchas that wasted my time:
Out comes a single nova.onnx file. That's the whole model.
With the model in place, it still wouldn't fire β and this one had nothing to do with training.
openWakeWord expects to be fed audio in 1280-sample windows (80 ms at 16 kHz). My audio capture
was handing it 480-sample frames (30 ms), because that's the frame size the voice-activity
detector wanted. Fed the wrong window size, the detector returned scores near 0 on every frame
and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling
predict():
buffer.extend(frame)
while len(buffer) >= 1280:
window, buffer = buffer[:1280], buffer[1280:]
score = model.predict(window)
If your freshly trained model scores zero on everything, suspect the plumbing before you blame the
training.
Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool
yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of
negatives, and a sweep across thresholds that prints recall and false positives per threshold.
evaluate_wakeword --positives eval/positives --negatives eval/negatives \
--thresholds 0.2,0.3,0.35,0.4,0.5
Use at least ~30 varied clips β different distances, speeds, background noise. In my set, about
a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell
you the truth.
This also killed a tempting assumption: that the detection threshold is the lever for recall.
It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no
threshold saves you β in fact lowering it from 0.4 to 0.3 made things worse (more false
positives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.
The synthetic-only model learned a TTS "Nova", not my voice in my room through my microphone.
The biggest, most reliable win was adding real recordings of me saying the word and retraining.
I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,
speed and background noise across ~60 of them (and held ~10% back for evaluation). In the training
notebook, those real positives get upsampled and mixed in with the synthetic ones.
The result, measured on the same evaluation set:
Recall went from 53% β 70% at threshold 0.35, just by folding in real-voice positives.
Not magic, but a real, measured step up β and exactly the kind of improvement that's invisible if
you're only judging by ear.
Here's the minimal path with public tools, so you can train your own keyword. Swap "nova" for
yours throughout.
1. Install and synthesize positives. Generate a few thousand clips, varying voice and speed:
pip install piper-tts openwakeword
python -m piper.download_voices --download-dir voices \
es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
python
import random, subprocess, pathlib
pathlib.Path("positives").mkdir(exist_ok=True)
voices = ["es_ES-sharvard-medium", "es_MX-ald-medium", "es_MX-claude-high"]
for i in range(2000):
v = random.choice(voices)
length = round(random.uniform(0.9, 1.3), 2) # speed variation
subprocess.run(
["piper", "--model", f"voices/{v}.onnx",
"--length-scale", str(length), "--output_file", f"positives/{v}_{i}.wav"],
input=b"nova\n",
)
Then listen to a handful before going further.
2. Train. Open openWakeWord's automatic_model_training.ipynb
(in their repo) on Kaggle or Colab, select a T4
GPU, set target_phrase = ["nova"], point it at your positives/ folder, and Run All. It pulls
the negative/background datasets and augmentation for you and exports a single nova.onnx.
3. Evaluate against your own clip folders and sweep thresholds β this is the step people skip:
import os, numpy as np, soundfile as sf
from openwakeword.model import Model
model = Model(wakeword_models=["nova.onnx"])
def best_score(path):
model.reset()
audio, _ = sf.read(path) # 16 kHz mono
audio = (audio * 32767).astype(np.int16)
top = 0.0
for i in range(0, len(audio) - 1280, 1280): # 80 ms windows
top = max(top, model.predict(audio[i:i + 1280])["nova"])
return top
pos = [best_score(f"eval/positives/{f}") for f in os.listdir("eval/positives")]
neg = [best_score(f"eval/negatives/{f}") for f in os.listdir("eval/negatives")]
for t in (0.2, 0.3, 0.35, 0.4, 0.5):
recall = sum(s >= t for s in pos) / len(pos)
fp = sum(s >= t for s in neg)
print(f"thr={t}: recall={recall:.0%} false_positives={fp}/{len(neg)}")
4. Fine-tune with real voice (optional, high impact). Record ~60 clips of yourself saying the
word (16 kHz mono β a few lines with sounddevice), hold back ~10% for evaluation, and add them to
the training set with upsampling. This is what took me from 53% to 70%.
5. Deploy. Copy nova.onnx to your device and feed the detector 1280-sample windows (see
the buffering snippet above). Tune the threshold from your evaluation numbers.
If you only remember four things:
And one bonus, because it bit me hardest: if a model scores zero on everything, check the window
size you're feeding it before you retrain anything.
Have you trained a custom wake word? I'd love to hear what moved recall for you β especially for
non-English keywords.
Want more context, or to see how we did it?
Nova isn't fully public yet, but you can get early access to the repository β all the
documentation and the complete source code β at https://gitlab.com/gabrielhruiz1/nova.
Leave us a message and we'll try to grant you access ASAP, until we publish everything
officially (we're still working on a few parts).