# Whistle: Speech to Text in 16.9 MB by Cactus Compute

> Source: <https://cactuscompute.com/blog/whistle>
> Published: 2026-10-02 21:22:43+00:00

Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as [Needle](https://cactuscompute.com/needle), from the same container and the same quantisation.

Whistle does three jobs, all of them on the device:

- **Transcription.** 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it.
- **Word timestamps.** Every word with its start, end and probability, aligned from the decoder's attention.
- **Speech embedding.** The encoder output, one row per 80 ms frame, without decoding a transcript.

## [The model](#the-model)

**The front end.** 16 kHz mono audio is framed at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames. A convolutional stem of 128 channels and kernel 9 halves that count three times, leaving 375 frames at one per 80 ms. Every stage after this runs at that rate, and `embed` returns one row per frame.

**The encoder.** Eight [Simple Attention](https://cactuscompute.com/blog/simple-attention-networks) blocks: four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. The attention is not causal. A frame at 3 s attends to a frame at 12 s.

**The decoder.** Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle's block list with a different layer count.

The speech-specific part is one addition per layer. Each decoder layer reads the encoder through a gated cross attention, `x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V`, with a gate learned per layer and K and V taken from the clip. Those projections run once when the clip arrives, 375 frames across 8 layers, and are then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.

**Decoding.** Five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over the phrases you pass in, alongside the beams, and lifts their log probability as the automaton advances. The transcript is capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language is emitted as a token rather than returned out of band.

**The ladder is on the decoder.** Every depth from 2 layers up was trained as a model of its own, and `--audio-depth` selects one at load time. The encoder is never sliced: all eight blocks run at every depth.

**Silence.** The engine measures the clip's loudness range before the decoder starts. Below the threshold it returns an empty transcript and an empty language, and never enters the beam search.

## [Benchmarks](#benchmarks)

Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9.

Each model ran on its official runtime at its defaults: Whistle's C++ engine at 5 beams, `openai-whisper`, and `moonshine-voice` non-streaming over whole audio. Time to first token is audio in to first token. Decode is tokens divided by the wall time after it, so the encoder is not counted twice. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.

Word error rates are scored with the Whisper normalizers. Whistle's are measured over 86,174 utterances. Whisper's and Moonshine's are the figures their authors published, from the multilingual checkpoints rather than the English-only ones. No test audio appears in Whistle's training or validation data, verified by comparing audio checksums and speaker IDs across every reported test set.

## [One engine, three ways to load it](#one-engine-three-ways-to-load-it)

`needle_load` reads whichever model a `.cact` file holds, so the same binary does speech, text, or both:

On the third line `needle_complete` takes the clip directly. The engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed `audio_`. No transcript is handled by the caller.

## [Get started](#get-started)

A 16 kHz WAV or raw samples need nothing beyond the base install. Other sample rates and microphone capture need the `[mic]` extra, which adds `soxr` and `sounddevice`.

Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second after it. `word_timestamps=True` adds each word with its times and probability. `keywords=["Siobhan", "Krzysztof"]` raises the log probability of those phrases during the search. `language="de"` forces the language instead of detecting it. `needle.Whistle()` is the same model as an object, for `embed(audio)` or to hold one tuned `.cact`.

`needle whistle playground` transcribes from the microphone in the terminal, and `needle whistle compare` runs the same clip through Whistle, Whisper and Moonshine side by side with their timings.

## [Deploy](#deploy)

The engine ships prebuilt for seventeen targets, from macOS and Linux through Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. Every folder holds a `needle` binary, `libneedle.a` and `needle.h`, and loads any `.cact` you hand it.

`needle_load`, `needle_transcribe` and `needle_embed` are the whole speech C API. The engine reads no environment variables. Every behaviour is a compiled default or an explicit flag.

Weights are on [Hugging Face](https://huggingface.co/Cactus-Compute/whistle), the engine and its platform folders are in [Cactus-Compute/needle3](https://huggingface.co/Cactus-Compute/needle3), and the source is on [GitHub](https://github.com/cactus-compute/needle).
