Cactus Compute releases a 16.9MB speech model for local CPUs Cactus Compute released Whistle on October 2nd, a 16.9MB open speech-to-text model that runs locally on a CPU inside the company's existing Needle runtime. Co-founder and CTO Henry Ndubuaku led the work, which supports English, German, French, Spanish, Italian, Dutch and Polish, transcribes up to 30 seconds of 16 kHz mono audio in one pass, and returns word-level timestamps and probabilities plus keyword biasing. The benchmarks are company-reported, and Cactus Compute's launch material does not establish how reliably the features perform across real users, accents and noisy environments. Cactus Compute releases a 16.9MB speech model for local CPUs Co-founder Henry Ndubuaku helped build Whistle into Cactus Compute's existing Needle runtime, adding keyword biasing and word timestamps for on-device applications. By RuntimeWire Staff https://runtimewire.com/author/runtimewire-staff ยท Published Primary source: Cactus https://www.cactuscompute.com/blog/whistle Why it matters Whistle extends Cactus's on-device strategy from local tool-calling models into speech. Its benchmarks are company-reported, and performance still needs to be judged on the hardware and audio that real applications use. Cactus Compute https://cactuscompute.com/?ref=runtimewire released Whistle https://runtimewire.com/models/huggingface/cactus-compute-whistle-3da55c53ab08c572 on October 2nd, an open speech-to-text model designed to run locally on a CPU. At 16.9MB, it is built to fit the same Needle runtime https://github.com/cactus-compute/needle?ref=runtimewire that Cactus Compute uses for small on-device language models, extending co-founder Henry Ndubuaku https://x.com/Henry Ndubuaku?ref=runtimewire 's work on compact inference into speech recognition. The release is covered in Cactus Compute's technical launch post https://www.cactuscompute.com/blog/whistle?ref=runtimewire . Cactus Compute's release thread on X https://x.com/cactuscompute/status/2106083041265562075?ref=runtimewire Ndubuaku is Cactus Compute's co-founder and CTO. His work has spanned fundamental AI research, distributed deep learning, GPU-kernel engineering and inference on small devices; before Cactus Compute, he was a research engineer on mobile AI models at an Imperial College spinout, according to his Forbes Technology Council profile https://councils.forbes.com/profile/Henry-Ndubuaku-Co-Founder-CTO-Cactus-Compute-Inc/d057a409-604d-4e5b-a8f2-a5bf2dce84b1?ref=runtimewire . Cactus Compute CEO Roman Shemet https://x.com/RomanShemet?ref=runtimewire , a former quant and economist with product and data-engineering experience, met Ndubuaku through Y Combinator's co-founder-matching program in London, according to YC's company profile https://www.ycombinator.com/companies/cactus-compute?ref=runtimewire . Their bet is that useful AI can move onto the devices people already carry, rather than requiring every interaction to travel to a server. A speech model built to fit the existing stack Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. Cactus Compute says it can transcribe a 16 kHz mono recording of up to 30 seconds in one pass, return word-level timestamps and probabilities, and provide speech embeddings without generating a transcript. Developers can also pass a list of keywords to bias recognition toward names or domain-specific terms. Cactus Compute says low-volume silence and steady noise return an empty transcript rather than a guessed sentence. Keyword biasing could help an application catch a customer's name; word timestamps can support highlighting, seeking or editing; and an empty result on silence addresses a familiar failure mode in voice interfaces. The launch material does not establish how reliably these features perform across real users, accents and noisy environments. The model uses a log-mel audio front end, a convolutional stem, an eight-block audio encoder and a decoder with gated cross-attention. Cactus Compute says the decoder can be loaded at different depths, while the audio encoder stays at full depth. Whistle also shares Needle https://runtimewire.com/models/huggingface/cactus-compute-needle-f0237a0d9e0908e4 's .cact model container and CPU inference engine. In practice, Cactus Compute is adding speech input to a deployment system it already wants developers to use for local tool-calling models. This builds on Cactus Compute's earlier Needle 2 release https://runtimewire.com/article/cactus-needle-2-14mb-agent-model-tiny-devices , which focused on running tool-calling models on low-cost devices. Whistle can transcribe audio in the same engine, and Cactus Compute says a combined setup can pass a clip directly to a language model and return structured tool calls. That makes voice a route into local device actions as well as a transcription feature. Cactus Compute's benchmark results Cactus Compute reports Whistle at 4.31% word error rate on LibriSpeech test-clean and 10.49% on test-other, compared with 4.9% and 11.0% for Whisper base https://runtimewire.com/models/huggingface/openai-whisper-base-6a4aa6291c2ec40a . It reports a 21.4 average on FLEURS versus 24.5 for Whisper base. Cactus Compute also lists results on SPGISpeech and Earnings-22, where it says Whistle scores 7.65 and 19.01 respectively. Those results do not show Whistle winning across the board. Cactus Compute's own benchmark page says Whisper base performs better on TED-LIUM, AMI and the MLS average. The comparisons use published results for Whisper and Moonshine where available, and Cactus Compute says the models ran on their official runtimes with default settings. It reports that its word-error-rate tests cover 86,174 utterances and that it checked the test sets against Whistle's training and validation data. Those details describe the methodology, but the figures remain company-reported rather than an independent reproduction. On an Apple M4 Pro CPU, Cactus Compute reports 11.1 milliseconds to first token and 1,319 decoded tokens per second for Whistle, against 73.2 milliseconds and 266 tokens per second for Whisper base. The speed comparison uses ten seconds of audio and separates time to first token from subsequent decoding. Cactus Compute also says its engine ships for 17 platform targets, ranging from desktop and mobile systems to RISC-V, MIPS, WebAssembly and WASI. The reported speed is tied to the M4 Pro test; it does not establish equivalent performance on every listed device. Why the deployment choice matters Whistle is aimed at applications where sending audio to a cloud service is undesirable or unavailable: wearables, phones, robots, cars and smart-home devices. Local inference can avoid a network round trip and keep audio on the device, while a compact model can be easier to fit into constrained hardware. Cactus Compute has also described a hybrid approach in which local inference handles routine audio and cloud processing can take more difficult segments. Whistle's local-only design and the broader hybrid product are distinct deployment choices. Whistle's scope is limited: it supports seven languages and caps a single pass at 30 seconds; the published comparisons show wins over Whisper base on some datasets and losses on others. Teams evaluating it will need to test their own audio, target hardware and vocabulary. Cactus Compute includes a compare command that runs a clip through Whistle, Whisper and Moonshine with timing results, and its launch materials provide a Python install path through cactus-needle . For Ndubuaku and Shemet, the release extends their engineering thesis: make local models small enough to run on ordinary devices, then build useful product behavior around them. Whistle adds speech to that stack. The next test is whether developers can get dependable transcription on the constrained devices Cactus Compute targets, not just strong scores on selected benchmarks.