Whistle: 16.9 MB Speech Model Runs AI Agents On-Device Cactus Compute released Whistle on October 2, a 16.9 MB on-device speech recognition model that runs entirely on CPU with zero dependencies and, paired with Cactus's Needle runtime, converts audio directly into structured tool calls without a server. Whistle posts a 4.31% word error rate on LibriSpeech test-clean versus Whisper base's 4.90%, 11.1 ms first-token latency versus Whisper base's 73.2 ms, and 1,319 tokens per second versus Whisper's 266 tok/s, while fitting in microcontroller RAM budgets that Whisper's 145.3 MB cannot. The model supports seven languages and 17 platform targets including RISC-V, WebAssembly and WASI, but lags Whisper base on TED-LIUM (7.61% vs 5.00%) and AMI, and carries a 30-second single-pass cap. Cactus Compute released Whistle on October 2 — a 16.9 MB on-device speech recognition model that runs entirely on CPU with zero dependencies. To put that number in context: eight copies of Whistle fit inside a single Whisper base model 145.3 MB . First-token latency is 11.1 milliseconds, six times faster than Whisper base’s 73.2 ms. Paired with Cactus’s Needle runtime https://huggingface.co/Cactus-Compute/whistle , Whistle does something no other small speech model does: it converts audio to structured tool calls in one binary, on-device, without touching a server. Whistle vs Whisper: The On-Device Benchmark Case The benchmark table tells a clean story. Whistle posts 4.31% word error rate on LibriSpeech test-clean — better than Whisper base’s 4.90% and Moonshine tiny v2’s 4.52%. On decode speed, Whistle hits 1,319 tokens per second, roughly five times Whisper’s 266 tok/s. The size advantage compounds the more constrained the target device: 16.9 MB fits comfortably in microcontroller RAM budgets where Whisper’s 145 MB simply cannot go. According to independent benchmarks from InsideTheLoop https://insidetheloop.dev/posts/cactus-whistle-on-device-speech-to-text , the results hold across multiple hardware targets. However, Whistle is not a universal replacement for Whisper. It lags on TED-LIUM 7.61% vs Whisper base’s 5.00% and AMI — conversational and meeting audio where Whisper’s larger encoder has an edge. The 30-second single-pass hard cap means longer recordings need chunking. Seven languages are supported English, German, French, Spanish, Italian, Dutch, Polish , which will block global deployments outside those markets. These are deliberate trade-offs for a model targeting embedded devices, not podcast transcription. Voice to Action, No Cloud Required The architecture story is what separates Whistle from every other small speech model on the market. Cactus designed it to run inside the same C++ engine as Needle, their on-device LLM runtime. The two share memory without IPC overhead. A single command takes audio in and returns structured function calls out: needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav Returns: {"function": "lights.turn off", "room": "kitchen"} That replaces a pipeline that previously required a separate Whisper.cpp transcription step, a string passed to an LLM API call, JSON parsing, then action execution. For on-device agents — the kind running in a wearable, a robot, or a home hub — eliminating that chain is significant. Latency drops. Complexity drops. Nothing leaves the device. Additionally, keyword biasing compounds the utility in real deployments. Developers can supply a list of domain-specific terms — product names, room names, command vocabulary — dropping biased word error rate from 18.43% to 4.46% with just 100 keywords. For agents with a bounded command vocabulary, that accuracy delta matters far more than generic benchmark WER. Related: DwarfStar 4: Run DeepSeek V4 Flash Locally at 39 t/s https://byteiota.com/dwarfstar-4-run-deepseek-v4-flash-locally-at-39-t-s/ 17 Platforms: Where Whistle Runs Cactus lists 17 platform targets: macOS Apple Silicon and Intel , Linux x86 64 and ARM64, Android, iOS, Windows ARM, RISC-V, WebAssembly, and WASI. The RISC-V and WASM targets are the ones that matter most — they cover ESP32-class microcontrollers and browser-side inference, environments where no other competitive speech model currently reaches. Install is one command: pip install cactus-needle , or use the C API for embedded targets without a Python runtime. The Hacker News community reaction https://news.ycombinator.com/item?id=50019911 881 points on October 9 confirms this deployment breadth is what excited developers most. Furthermore, the model ships frame-level speech embeddings alongside transcription — one row per 80 ms of audio. This opens a secondary use case: voice fingerprinting and retrieval without storing raw audio, useful for on-device speaker identification or wake-word-style matching in the same pipeline. Key Takeaways - Whistle is 8.6x smaller than Whisper base and 6x faster to first token — the size and latency thresholds that make on-device speech practical for microcontrollers, wearables, and embedded agents - The Needle integration converts audio to structured tool calls in one binary, replacing a multi-step cloud pipeline with a single on-device command - Keyword biasing drops biased WER from 18.43% to 4.46% with 100 domain terms — critical for agents with a bounded command vocabulary - Trade-offs are real: 30-second single-pass cap, 7-language scope only, and weaker accuracy on conversational audio. Test against your actual environment before committing - 17 platform targets including RISC-V and WASM — the first speech model that meaningfully reaches microcontroller-class devices without a cloud dependency