Voxlocal: a minimal voice agent written in Rust Developer samkhawase released voxlocal, a fully local, low-latency voice agent for macOS written in Rust, on 9 Oct 2026. The CLI tool targets a narrow automotive-service use case in English and combines off-the-shelf Whisper speech recognition and Piper text-to-speech with a turn pipeline that normalizes text, retrieves context via cosine similarity, routes requests to tools, and generates replies, deliberately omitting async streaming and voice activity detection to keep each stage inspectable. The author warns it is an experimental project and not for production use. Voxlocal: a minimal voice agent written in Rust 9 Oct 2026 16 min read “ man muss immer umkehren” – Carl Gustav Jacob Jacobi “When in doubt, subtract” I’ve been working with voice agents for a while now, and they’re a fascinating piece of technology. It is magical to talk to a voice agent and get the work done. Ever wondered how voice agents work behind the scenes? This blog post is my attempt to understand their inner workings. What is even a Voice agent? An AI agent is an autonomous software system that uses AI models to reason, plan steps, and execute tasks to reach a specific goal. Voice agent is a version of it — A system that can engage in natural human-like spoken conversations to complete tasks. In simple words, voice agents listen, think, do tasks, and speak. Bird’s eye view That’s a mouthful of a diagram. Since I was curious how all of this works, I created a bare-bones project in Rust to understand the finer details of what’s happening behind the scenes. Let us get the legacy components out of our way before we dive into the voice agents part. Telephony Provider Telephony networking is a complex topic and way beyond the scope of this article. In our case, the Telephony Provider segment bridges legacy public telephone networks PSTN/SIP with cloud-based AI applications. It does complex tasks like handling the call lifecycle, bi-directional streaming, transcoding and keeping the latency low. One major task this segment does is jitter buffering for dropped or delayed packets, and audio upsampling to preserve STT accuracy. This segment tries to keep the network latency under the threshold. Modern telephony is an engineering marvel and closely related to the modern internet. With that out of the way, let’s take a look at what we are cooking. Introducing voxlocal, our avant-garde voice agent Voxlocal is a fully local, low-latency voice agent for macOS. It’s a CLI tool written in Rust that lets the user talk to the voice agent and get a response back. In the spirit of experimentation, I have used bare bones components and micro models to get as close to the magic as possible. Don’t use it in production Its focus is very narrow: An automotive-service use case, through English language. I’m also using a few off-the-shelf components like Whisper and Piper instead of reinventing the wheel. To keep the pipeline inspectable and transparent, voxlocal deliberately avoids complex async streaming or VAD Voice Activity Detection . I might dive deeper into them in the upcoming revisions. Explore the source code: Open-source project · Rust · macOS samkhawase / voxlocal A minimal, fully local voice agent built to make each step of the conversation easy to inspect. View repository https://github.com/samkhawase/voxlocal Let us explore it step by step 1. What happens when you speak to a voice agent? When you talk to a voice agent, or give it an audio sample, it breaks it down into multiple smaller tasks to do the work. It recognizes the words from the audio and cleans up the text by removing ‘Um’s, ‘ah’s. It uses the words to find relevant information and chooses an action based on them. Finally it turns the response back into speech. In voxlocal , microphone audio is captured and passed to speech recognition first: js let capture = record until RECORD SECS ?; let transcript = stt.transcribe &capture.samples ?; The transcript then enters the shared turn pipeline. In the pipeline the agent normalizes the text, searches for matching context, routes the request to the relevant tool, and finally prepares a reply: run turn &norm, &rag, &embedder, &mut llm, tts.as mut , transcript, &mut budget, ?; Each of the stages use a different technique to carry out their operations. - Speech recognition determines the words from the audio - Retrieval compares the text blocks in its numerical form using Maths