cd /news/ai-tools/training-a-wake-word-that-actually-f… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-147510] src=dev.to β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Training a wake word that actually fires: lessons from an offline voice assistant

A developer built an offline wake-word detector for a Raspberry Pi 4 voice assistant named "Nova" using openWakeWord and Piper, finding that training positives with Spanish TTS voices instead of English was the single biggest recall improvement. The engineer also traced a total detection failure to a window-size mismatch: openWakeWord expects 1280-sample (80 ms) windows at 16 kHz, while the capture pipeline was feeding 480-sample (30 ms) frames, yielding scores near zero until frames were buffered.

by read7 min views5 publishedOct 8, 2026

I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.

The very first thing it has to do is also the easiest to underestimate: notice when you say its

name. If the wake word doesn't fire, nothing else in the pipeline ever gets a turn β€” the

speech-to-text, the LLM, the voice, all of it sits there waiting.

My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what

actually moved the needle was not where I expected. This post is the recipe I wish I'd had β€” with

the specific mistakes that cost me recall, so you can skip them.

Everything here uses open-source tools: openWakeWord

for the model and Piper for synthetic speech. The approach

works for any keyword, not just mine.

openWakeWord doesn't train a giant speech model. It trains a small classifier on top of a

shared, pre-trained audio embedding. The pipeline is:

audio β†’ melspectrogram β†’ shared embedding β†’ small "is this the wake word?" classifier

That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the

tiny classifier is yours. You feed it two things: positives (lots of people saying your word)

and negatives (speech and noise that is not your word). The catch is that you rarely have

thousands of real recordings of a made-up name β€” so you synthesize them.

The standard trick is to generate thousands of utterances of your wake word with a text-to-speech

engine, varying voice, speed and pitch. I used Piper.

Here is the lesson, and it's the big one: the language of the TTS voice has to match how you'll actually say the word. My first model was trained with English voices. In English, "Nova" is

Switching the positives to Spanish Piper voices was the single biggest improvement I made.

After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly

(es_ES-sharvard-medium, es_MX-ald-medium, es_MX-claude-high) and dropped the ones that sounded

off on this particular word.

python -m piper.download_voices --download-dir voices \
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high

Do this before you train anything: generate ~50 clips and actually listen. One minute of

listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you

hear in those clips is what the model is about to learn.

Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's

training uses large public datasets of general speech and noise as negatives, plus pre-computed

features to validate false positives. On top of that it augments the positives by mixing in

noise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.

You mostly get this for free from the project's training notebook β€” just don't skip it.

Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google

Colab and Kaggle both work and are free. The config change that makes it your word is a single

line:

target_phrase = ["nova"]

Two hardware gotchas that wasted my time:

Out comes a single nova.onnx file. That's the whole model.

With the model in place, it still wouldn't fire β€” and this one had nothing to do with training.

openWakeWord expects to be fed audio in 1280-sample windows (80 ms at 16 kHz). My audio capture

was handing it 480-sample frames (30 ms), because that's the frame size the voice-activity

detector wanted. Fed the wrong window size, the detector returned scores near 0 on every frame

and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling

predict():

buffer.extend(frame)
while len(buffer) >= 1280:
    window, buffer = buffer[:1280], buffer[1280:]
    score = model.predict(window)

If your freshly trained model scores zero on everything, suspect the plumbing before you blame the

training.

Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool

yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of

negatives, and a sweep across thresholds that prints recall and false positives per threshold.

evaluate_wakeword --positives eval/positives --negatives eval/negatives \
    --thresholds 0.2,0.3,0.35,0.4,0.5

Use at least ~30 varied clips β€” different distances, speeds, background noise. In my set, about

a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell

you the truth.

This also killed a tempting assumption: that the detection threshold is the lever for recall.

It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no

threshold saves you β€” in fact lowering it from 0.4 to 0.3 made things worse (more false

positives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.

The synthetic-only model learned a TTS "Nova", not my voice in my room through my microphone.

The biggest, most reliable win was adding real recordings of me saying the word and retraining.

I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,

speed and background noise across ~60 of them (and held ~10% back for evaluation). In the training

notebook, those real positives get upsampled and mixed in with the synthetic ones.

The result, measured on the same evaluation set:

Recall went from 53% β†’ 70% at threshold 0.35, just by folding in real-voice positives.

Not magic, but a real, measured step up β€” and exactly the kind of improvement that's invisible if

you're only judging by ear.

Here's the minimal path with public tools, so you can train your own keyword. Swap "nova" for

yours throughout.

1. Install and synthesize positives. Generate a few thousand clips, varying voice and speed:

pip install piper-tts openwakeword
python -m piper.download_voices --download-dir voices \
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
python
import random, subprocess, pathlib
pathlib.Path("positives").mkdir(exist_ok=True)
voices = ["es_ES-sharvard-medium", "es_MX-ald-medium", "es_MX-claude-high"]
for i in range(2000):
    v = random.choice(voices)
    length = round(random.uniform(0.9, 1.3), 2)   # speed variation
    subprocess.run(
        ["piper", "--model", f"voices/{v}.onnx",
         "--length-scale", str(length), "--output_file", f"positives/{v}_{i}.wav"],
        input=b"nova\n",
    )

Then listen to a handful before going further.

2. Train. Open openWakeWord's automatic_model_training.ipynb

(in their repo) on Kaggle or Colab, select a T4

GPU, set target_phrase = ["nova"], point it at your positives/ folder, and Run All. It pulls

the negative/background datasets and augmentation for you and exports a single nova.onnx.

3. Evaluate against your own clip folders and sweep thresholds β€” this is the step people skip:

import os, numpy as np, soundfile as sf
from openwakeword.model import Model

model = Model(wakeword_models=["nova.onnx"])

def best_score(path):
    model.reset()
    audio, _ = sf.read(path)                       # 16 kHz mono
    audio = (audio * 32767).astype(np.int16)
    top = 0.0
    for i in range(0, len(audio) - 1280, 1280):    # 80 ms windows
        top = max(top, model.predict(audio[i:i + 1280])["nova"])
    return top

pos = [best_score(f"eval/positives/{f}") for f in os.listdir("eval/positives")]
neg = [best_score(f"eval/negatives/{f}") for f in os.listdir("eval/negatives")]
for t in (0.2, 0.3, 0.35, 0.4, 0.5):
    recall = sum(s >= t for s in pos) / len(pos)
    fp = sum(s >= t for s in neg)
    print(f"thr={t}: recall={recall:.0%}  false_positives={fp}/{len(neg)}")

4. Fine-tune with real voice (optional, high impact). Record ~60 clips of yourself saying the

word (16 kHz mono β€” a few lines with sounddevice), hold back ~10% for evaluation, and add them to

the training set with upsampling. This is what took me from 53% to 70%.

5. Deploy. Copy nova.onnx to your device and feed the detector 1280-sample windows (see

the buffering snippet above). Tune the threshold from your evaluation numbers.

If you only remember four things:

And one bonus, because it bit me hardest: if a model scores zero on everything, check the window

size you're feeding it before you retrain anything.

Have you trained a custom wake word? I'd love to hear what moved recall for you β€” especially for

non-English keywords.

Want more context, or to see how we did it?

Nova isn't fully public yet, but you can get early access to the repository β€” all the

documentation and the complete source code β€” at https://gitlab.com/gabrielhruiz1/nova.

Leave us a message and we'll try to grant you access ASAP, until we publish everything

officially (we're still working on a few parts).

── more in #ai-tools 4 stories Β· sorted by recency
── more on @openwakeword 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/training-a-wake-word…] indexed:0 read:7min 2026-10-08 Β· β€”