# Training a wake word that actually fires: lessons from an offline voice assistant

> Source: <https://dev.to/gabrielhruiz/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant-3716>
> Published: 2026-10-08 11:14:26+00:00

I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.

The very first thing it has to do is also the easiest to underestimate: notice when you say its

name. If the **wake word** doesn't fire, nothing else in the pipeline ever gets a turn — the

speech-to-text, the LLM, the voice, all of it sits there waiting.

My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what

actually moved the needle was *not* where I expected. This post is the recipe I wish I'd had — with

the specific mistakes that cost me recall, so you can skip them.

Everything here uses open-source tools: [openWakeWord](https://github.com/dscripka/openWakeWord)

for the model and [Piper](https://github.com/rhasspy/piper) for synthetic speech. The approach

works for any keyword, not just mine.

openWakeWord doesn't train a giant speech model. It trains a **small classifier** on top of a

shared, pre-trained audio embedding. The pipeline is:

```
audio → melspectrogram → shared embedding → small "is this the wake word?" classifier
```

That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the

tiny classifier is yours. You feed it two things: **positives** (lots of people saying your word)

and **negatives** (speech and noise that is *not* your word). The catch is that you rarely have

thousands of real recordings of a made-up name — so you synthesize them.

The standard trick is to generate thousands of utterances of your wake word with a text-to-speech

engine, varying voice, speed and pitch. I used Piper.

Here is the lesson, and it's the big one: **the language of the TTS voice has to match how you'll actually say the word.** My first model was trained with English voices. In English, "Nova" is

Switching the positives to **Spanish Piper voices** was the single biggest improvement I made.

After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly

(`es_ES-sharvard-medium`, `es_MX-ald-medium`, `es_MX-claude-high`) and dropped the ones that sounded

off on this particular word.

```
# Grab a few Spanish voices and generate a small batch to listen to FIRST
python -m piper.download_voices --download-dir voices \
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
# then synthesize many "nova" clips at 16 kHz mono, varying voice/speed/pitch
```

**Do this before you train anything:** generate ~50 clips and actually *listen*. One minute of

listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you

hear in those clips is what the model is about to learn.

Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's

training uses large public datasets of general speech and noise as negatives, plus pre-computed

features to validate false positives. On top of that it **augments** the positives by mixing in

noise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.

You mostly get this for free from the project's training notebook — just don't skip it.

Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google

Colab and Kaggle both work and are free. The config change that makes it *your* word is a single

line:

```
target_phrase = ["nova"]
```

Two hardware gotchas that wasted my time:

Out comes a single `nova.onnx` file. That's the whole model.

With the model in place, it still wouldn't fire — and this one had nothing to do with training.

openWakeWord expects to be fed audio in **1280-sample windows (80 ms at 16 kHz)**. My audio capture

was handing it **480-sample frames (30 ms)**, because that's the frame size the voice-activity

detector wanted. Fed the wrong window size, the detector returned scores near **0** on every frame

and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling

`predict()`:

```
# accumulate frames until we have a full 1280-sample window, then score
buffer.extend(frame)
while len(buffer) >= 1280:
    window, buffer = buffer[:1280], buffer[1280:]
    score = model.predict(window)
```

If your freshly trained model scores zero on *everything*, suspect the plumbing before you blame the

training.

Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool

yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of

negatives, and a sweep across thresholds that prints recall and false positives per threshold.

```
evaluate_wakeword --positives eval/positives --negatives eval/negatives \
    --thresholds 0.2,0.3,0.35,0.4,0.5
```

Use **at least ~30 varied clips** — different distances, speeds, background noise. In my set, about

a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell

you the truth.

This also killed a tempting assumption: that the **detection threshold** is the lever for recall.

It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no

threshold saves you — in fact lowering it from 0.4 to 0.3 made things *worse* (more false

positives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.

The synthetic-only model learned a TTS "Nova", not *my* voice in *my* room through *my* microphone.

The biggest, most reliable win was adding **real recordings** of me saying the word and retraining.

I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,

speed and background noise across ~60 of them (and held ~10% back for evaluation). In the training

notebook, those real positives get **upsampled** and mixed in with the synthetic ones.

The result, measured on the same evaluation set:

**Recall went from 53% → 70% at threshold 0.35**, just by folding in real-voice positives.

Not magic, but a real, measured step up — and exactly the kind of improvement that's invisible if

you're only judging by ear.

Here's the minimal path with public tools, so you can train your own keyword. Swap `"nova"` for

yours throughout.

**1. Install and synthesize positives.** Generate a few thousand clips, varying voice and speed:

```
pip install piper-tts openwakeword
python -m piper.download_voices --download-dir voices \
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
python
import random, subprocess, pathlib
pathlib.Path("positives").mkdir(exist_ok=True)
voices = ["es_ES-sharvard-medium", "es_MX-ald-medium", "es_MX-claude-high"]
for i in range(2000):
    v = random.choice(voices)
    length = round(random.uniform(0.9, 1.3), 2)   # speed variation
    subprocess.run(
        ["piper", "--model", f"voices/{v}.onnx",
         "--length-scale", str(length), "--output_file", f"positives/{v}_{i}.wav"],
        input=b"nova\n",
    )
# (CLI flags vary slightly by Piper version — check `piper --help`.)
```

Then **listen to a handful** before going further.

**2. Train.** Open openWakeWord's `automatic_model_training.ipynb`

([in their repo](https://github.com/dscripka/openWakeWord)) on Kaggle or Colab, select a **T4**

GPU, set `target_phrase = ["nova"]`, point it at your `positives/` folder, and *Run All*. It pulls

the negative/background datasets and augmentation for you and exports a single `nova.onnx`.

**3. Evaluate** against your own clip folders and sweep thresholds — this is the step people skip:

``` python
import os, numpy as np, soundfile as sf
from openwakeword.model import Model

model = Model(wakeword_models=["nova.onnx"])

def best_score(path):
    model.reset()
    audio, _ = sf.read(path)                       # 16 kHz mono
    audio = (audio * 32767).astype(np.int16)
    top = 0.0
    for i in range(0, len(audio) - 1280, 1280):    # 80 ms windows
        top = max(top, model.predict(audio[i:i + 1280])["nova"])
    return top

pos = [best_score(f"eval/positives/{f}") for f in os.listdir("eval/positives")]
neg = [best_score(f"eval/negatives/{f}") for f in os.listdir("eval/negatives")]
for t in (0.2, 0.3, 0.35, 0.4, 0.5):
    recall = sum(s >= t for s in pos) / len(pos)
    fp = sum(s >= t for s in neg)
    print(f"thr={t}: recall={recall:.0%}  false_positives={fp}/{len(neg)}")
```

**4. Fine-tune with real voice (optional, high impact).** Record ~60 clips of yourself saying the

word (16 kHz mono — a few lines with `sounddevice`), hold back ~10% for evaluation, and add them to

the training set with upsampling. This is what took me from 53% to 70%.

**5. Deploy.** Copy `nova.onnx` to your device and feed the detector **1280-sample windows** (see

the buffering snippet above). Tune the threshold from your evaluation numbers.

If you only remember four things:

And one bonus, because it bit me hardest: if a model scores zero on everything, check the window

size you're feeding it before you retrain anything.

Have you trained a custom wake word? I'd love to hear what moved recall for you — especially for

non-English keywords.

**Want more context, or to see how we did it?**

Nova isn't fully public yet, but you can get **early access** to the repository — all the

documentation and the complete source code — at **[https://gitlab.com/gabrielhruiz1/nova](https://gitlab.com/gabrielhruiz1/nova)**.

Leave us a message and we'll try to grant you access ASAP, until we publish everything

officially (we're still working on a few parts).
