cd /news/ai-agents/voxlocal-a-minimal-voice-agent-writt… · home › topics › ai-agents › article
[ARTICLE · art-148675] src=samkhawase.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Voxlocal: a minimal voice agent written in Rust

Developer samkhawase released voxlocal, a fully local, low-latency voice agent for macOS written in Rust, on 9 Oct 2026. The CLI tool targets a narrow automotive-service use case in English and combines off-the-shelf Whisper speech recognition and Piper text-to-speech with a turn pipeline that normalizes text, retrieves context via cosine similarity, routes requests to tools, and generates replies, deliberately omitting async streaming and voice activity detection to keep each stage inspectable. The author warns it is an experimental project and not for production use.

read15 min views2 publishedOct 10, 2026

9 Oct 2026

16 min read

“man muss immer umkehren” – Carl Gustav Jacob Jacobi

“When in doubt, subtract”

I’ve been working with voice agents for a while now, and they’re a fascinating piece of technology. It is magical to talk to a voice agent and get the work done. Ever wondered how voice agents work behind the scenes? This blog post is my attempt to understand their inner workings.

What is even a Voice agent?

An AI agent is an autonomous software system that uses AI models to reason, plan steps, and execute tasks to reach a specific goal. Voice agent is a version of it — A system that can engage in natural human-like spoken conversations to complete tasks.

In simple words, voice agents listen, think, do tasks, and speak.

Bird’s eye view

That’s a mouthful of a diagram. Since I was curious how all of this works, I created a bare-bones project in Rust to understand the finer details of what’s happening behind the scenes.

Let us get the legacy components out of our way before we dive into the voice agents part.

Telephony Provider

Telephony networking is a complex topic and way beyond the scope of this article. In our case, the Telephony Provider segment bridges legacy public telephone networks (PSTN/SIP) with cloud-based AI applications. It does complex tasks like handling the call lifecycle, bi-directional streaming, transcoding and keeping the latency low.

One major task this segment does is jitter buffering for dropped or delayed packets, and audio upsampling to preserve STT accuracy. This segment tries to keep the network latency under the threshold.

Modern telephony is an engineering marvel and closely related to the modern internet.

With that out of the way, let’s take a look at what we are cooking.

Introducing voxlocal, our avant-garde voice agent

Voxlocal is a fully local, low-latency voice agent for macOS. It’s a CLI tool written in Rust that lets the user talk to the voice agent and get a response back. In the spirit of experimentation, I have used bare bones components and micro models to get as close to the magic as possible. Don’t use it in production!

Its focus is very narrow: An automotive-service use case, through English language. I’m also using a few off-the-shelf components like Whisper and Piper instead of reinventing the wheel. To keep the pipeline inspectable and transparent,voxlocal deliberately avoids complex async streaming or VAD (Voice Activity Detection). I might dive deeper into them in the upcoming revisions.

Explore the source code:

Open-source project · Rust · macOS samkhawase / voxlocal A minimal, fully local voice agent built to make each step of the conversation easy to inspect. View repository

Let us explore it step by step

1. What happens when you speak to a voice agent?

When you talk to a voice agent, or give it an audio sample, it breaks it down into multiple smaller tasks to do the work. It recognizes the words from the audio and cleans up the text by removing ‘Um’s, ‘ah’s. It uses the words to find relevant information and chooses an action based on them. Finally it turns the response back into speech.

In voxlocal, microphone audio is captured and passed to speech recognition first:

let capture = record_until(RECORD_SECS)?;
let transcript = stt.transcribe(&capture.samples)?;

The transcript then enters the shared turn pipeline. In the pipeline the agent normalizes the text, searches for matching context, routes the request to the relevant tool, and finally prepares a reply:

run_turn(
    &norm, &rag, &embedder, &mut llm, tts.as_mut(),
    transcript, &mut budget,
)?;

Each of the stages use a different technique to carry out their operations.

  • Speech recognition determines the words from the audio
  • Retrieval compares the text blocks in its numerical form using Maths<sup>1</sup> (cosine similarity)
  • An LLM predicts an action
  • Plain old Rust code does the tool calls and executes the action.
  • Speech synthesis generates audio from the final reply

However this is too dry and academic. Let’s unwrap the layers through a working example where the user asks the agent: “How much does an oil change cost?”

2. From sound wave to words

I was surprised to learn that the microphone captures sound by measuring air pressure continuously and turns the readings into audio samples. voxlocal then converts the captured audio to mono<sup>2</sup> at 16,000 samples per second and then passes it on to Whisper for further processing.

Whisper is our trained neural network and it is designed to look for patterns in the audio that resemble speech sounds. It then estimates if the sound makes a sequence of words. The Whisper creators have trained it and fine tuned it for this task, our program simply piggybacks on it and calls the model<sup>3</sup>.

let mut params =
    whisper_rs::FullParams::new(
        whisper_rs::SamplingStrategy::Greedy { best_of: 1 }
    );
params.set_language(Some("en"));
params.set_temperature(0.0);
state.full(params, audio)
    .map_err(|e| anyhow!("whisper decode failed: {e}"))?;

For this example we use greedy decoding instead of sampling to pick the most likely next token. Temperature 0 means it will just return the first match of highest likelihood like oil or change. We can do this because our context is very limited and we know our dialogues.

Once we get the response from Whisper in segments, it is the job of voxlocal to join them into one transcript.

let mut out = String::new();
for i in 0..state.full_n_segments() {
    if let Some(segment) = state.get_segment(i) {
        if let Ok(text) = segment.to_str_lossy() {
            out.push_str(&text);
        }
    }
}
let transcript = out.trim().to_string();

In our case the output is:

“What does an oil change cost?”

At this stage the output is an estimate made from the output signal. It could vary a bit from the actual spoken words. Analog to digital transcribing is always tricky due to various factors like noise, accents, cadence, timbre etc. Hence we need an extra stage to clean it up and make it more consistent for further processing.

3. Cleaning up recognized speech

Humans are not always consistent when speaking and often add filler words like “Um”, awkward s or peculiar pronunciation of numbers (“noyynTherTee” instead of “Nine Thirty”). voxlocal applies simple and predictable rules to clean up the text before feeding it to the next stages.

For example, the normalizer in voxlocal removes common filler phrases, trims whitespace, and joins the separators:

s = self.thousands.replace_all(&s, "$1$2").into_owned();
s = self.fillers.replace_all(&s, " ").into_owned();
s = self.space.replace_all(&s, " ").into_owned();

When we say “Um yeah, what does an oil change cost?”, it might become “What does an oil change cost”. The program simply uses existing known transformations and lets the next stages interpret the meaning behind it. It’s simply regular expressions which are reused on each turn, keeping in line with our bare bones approach.

4. Turning words into vectors

Now that voxlocal has got the words sanitized, it need to find users intent from it. User can say “What does an oil change cost” and “How much for an oil service” both have the same intent. So how do we figure out the intent?

The answer is to use an embedding model, which converts the text into a list of numbers (you can say vector to sound cool). With this trick text with related meanings tend to have vectors pointing in the same direction.

We use MiniLM in voxlocal to create one vector for the normalized query.

pub fn embed_one(&self, text: &str) -> Result<Vec<f32>> {
    Ok(self.embed_batch(&[text])?.remove(0))
}

Inside the function embed_batch, MiniLM produces a representation for each token. The code then combines these token representations into one sentence vector using masked mean pooling. In simple words: it averages the real tokens and ignores padding added to make a batch the same length.

let mask = attention.to_dtype(DType::F32)?.unsqueeze(2)?;
let summed = hidden.broadcast_mul(&mask)?.sum(1)?;
let counts = mask.sum(1)?;
let pooled = summed.broadcast_div(&counts)?.contiguous()?;

How does the model determine which patterns matter for meaning and intent? The answer is in the model’s learned weights, which are pretrained and ready to use for voxlocal. The pooling code of voxlocal then reduces the token-level outputs to a single vector for the complete sentence. This is purely the model’s output representation; the Rust code does not define the meaning of each coordinate.

Finally, the vector is L2-normalized.

let norm = row.iter().map(|v| v * v).sum::<f32>().sqrt();
let unit_vector = row.iter().map(|v| v / norm).collect::<Vec<_>>();

In simple words, normalization makes sure every nonzero vector has length one. This is needed for the next stage which compares direction using a simple dot product. According to Maths, similar meanings point in the similar direction even when the wordings differ. And yes, it literally calculates the cosine of the angle literally and geometrically, if you are wondering. I can see my high school maths teacher Mr. Bhalérao, smiling at me with the smug look.

5. Finding relevant knowledge

With the query which is a unit-length vector voxlocal is now ready to compare it with the existing service documents. voxlocal compares the query vector with each document vector and finds useful context.

The dot products of the vectors, which is cosine similarity, are calculated and compared for similarity. A better match means the vectors point in a more similar direction. Our search compared the scores and drops the ones that are below our reference threshold. The remaining matches are sorted and the request number is returned.

let mut scored: Vec<(&Doc, f32)> = self.docs
    .iter()
    .map(|doc| {
        let sim: f32 = doc.embedding.iter()
            .zip(query.iter())
            .map(|(a, b)| a * b)
            .sum();
        (doc, sim)
    })
    .filter(|(_, sim)| *sim >= RAG_MIN_SIM)
    .collect();
scored.sort_by(|a, b| {
    b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal)
});
scored.truncate(top_k);

We need the threshold because we want to discard documents that rank higher but are irrelevant.

If it’s getting too mathematical, let me explain in simple terms: Imagine the vector of our query is [1.0, 0.0], and a document vector pointing in a similar direction has a vector [0.8, 0.6]. The dot product of them is 1.0×0.8 + 0.0×0.6 = 0.8. A vector (unrelated term) pointing mostly another way and having a vector [0.1, 0.995], scores about 0.1. Whichever dot product is closer to 1.0 is the winner. The query match yields $\text{Similarity} = (1.0 \times 0.8) + (0.0 \times 0.6) = 0.8$.

A document must also score at least 0.35 to be included. The sample oil-change query scored 0.816 while an unrelated greeting such as “hi” scores around 0.10–0.15 against this corpus. This helps us weed out unrelated terms.

This is a minimal RAG example and the search returns only 2 documents at max. The pipeline prepares them for the language model and caps the context at 320 chars. The language model decides what response and action to perform based on this evidence.

6. How a language model chooses an action

We now have the user’s question and the retrieved context available for the next stage. In voxlocal, our language model (SmolLM2) is prompted to decide and return a structured tool call instead of a free-form answer.

E.g. For “What does an oil change cost?”, the intended output is a tool name plus its arguments:

{
  "tool": "check_price",
  "args": {
    "service": "oil change"
  }
}

The model produces this output one token at a time and a token can be a whole word, part of a word, punctuation, or a JSON fragment. At each step, our model predicts which token is the most likely the next one based on the preceding prompt and tokens. The Rust code runs that generation through Candle and it uses the model files which contain the learned information.

The router limits how many tokens can be generated, and then parses the result as a ToolCall:

let (raw, truncated) = self.generate(&prompt, max_new)?;
let call = parse_tool_call(&raw)
    .or_else(|| truncated.then(|| repair_truncated(&raw)).flatten())
    .ok_or_else(|| anyhow!("model did not emit valid tool JSON: {raw:?}"))?;

This snippet above shows the boundary between generation and interpretation. The model returns a text value, and the program attempts to interpret that text as a canonical action. The model’s output is not automatically treated as an executable command by voxlocal.

Small language models can run into trouble and stop mid-JSON or produce malformed output. When it notices a generation was truncated, voxlocal attempts a repair to overcome this. If the parsing still fails, the router component of voxlocal returns an error for that turn. The next section explains how a valid tool call is checked and executed by the application code.

7. From model output to action

As we’ve seen, the model-generated tool call is still text at this stage. In order to perform the specified action, voxlocal has to parse it and decide what the tool is allowed to do.

voxlocal parses the model’s JSON and then maps its proposed tool name onto the supported tool set. It also takes the service from the retrieved result in case one is available:

let call = parse_tool_call(&raw)
    .or_else(|| truncated.then(|| repair_truncated(&raw)).flatten())
    .ok_or_else(|| anyhow!("model did not emit valid tool JSON: {raw:?}"))?;
let mut call = call;
call.tool = snap_tool(&call.tool, query);
if let Some(hit) = top_hit {
    call.args.insert(
        "service".to_string(),
        serde_json::Value::String(hit.title.clone()),
    );
}

What we see here is an interesting boundary between a statistical prediction and programmatic behavior. The rust program parses the suggestions from the model and applies rules before execution.

For example, the router extracts a time from the current query instead of blindly trusting a time the model may have copied from an example.

The executor then handles the supported actions we have setup for this example:

match call.tool.as_str() {
    "book_appointment" => { /* prepare booking reply */ }
    "check_price" => { /* look up service price */ }
    other => format!("Unknown tool {other}."),
}

For this project nothing is actually executed, because this is a mock. In real life the app might call a calendar or an API.

8. Turning a reply back into sound

At this stage we are ready with an answer for the user and it’s time for the program to convert the text response back into speech. Piper converts words into an audio representation and then using its trained voice model generates the final audio. Like in the case of Whisper, we are piggybacking on the learned model weights that help us with timing, pitch, timbre etc.

voxlocal calls Piper through piper-rs and synthesize the voice:

pub fn speak(&mut self, text: &str) -> Result<(Vec<f32>, u32)> {
    self.piper
        .create(text, false, None, None, None, None)
        .map_err(|e| anyhow!("piper synth failed: {e}"))
}

voxlocal passes the resulting audio samples to rodio for playback, along with their sample rate.

let source = rodio::buffer::SamplesBuffer::new(
    1,
    sample_rate,
    samples.to_vec(),
);
player.append(source);
player.sleep_until_end();

We need to keep in mind that the neural synthesis happens inside Piper’s ONNX voice model<sup>4</sup>, not in this Rust wrapper. voxlocal gives Piper text and receives audio which is then played by rodio through the selected output device.

This final step completes the path from the caller’s speech to the agent’s spoken response. In a phone system, the final audio would be sent back through the call’s media stream instead of being played through local speakers.

Measuring the Pipeline #

9. Where the time goes

The voice agent’s response time includes both computation and time spent recording or playing audio. voxlocal has a five-second recording window, but the model computation is not 5 seconds. We can see the details by running the program locally and here is one (abridged) request that shows the pipeline. The logs at different stages show the transformations and the latency budget shows how long the measured stages took.

[Stage 1: Voice capture]: capture complete: 80554 samples (5150 ms)
[Stage 2: Speech to text]: produced 37 characters: "What is the price for the oil change?"
[Stage 3: Text normalization]: output: "What is the price for the oil change?" (0.9 ms)
[Stage 4: Retrieval augmented generation]: selected document: oil change (0.816)
[Stage 4: Retrieval augmented generation]: using 86 context characters (12.2 ms)
[Stage 5: LLM tool routing]: generating up to 48 tokens with 86 context characters
  [llm]    {"tool": "check_price", "args": {"service": "brake pads"}}
  [tool]   check_price {"service": String("oil change")}
[Stage 5: LLM tool routing]: reply prepared: "oil change is $79." (68.6 ms)
[Stage 6: Text to speech]: generated 47616 samples at 22050 Hz (72.1 ms)
[Stage 7: Audio playback]: playback complete
latency budget:
  capture     5149.6 ms
  stt           55.1 ms
  normalize      0.9 ms
  rag           12.2 ms
  llm           68.6 ms
  tts           72.1 ms
  play        1135.3 ms
  TOTAL       6493.9 ms  <- total compute

Note: Interestingly, the model mistakenly reused brake pads from the appointment example in its prompt. The router corrected the service to oil change using the retrieved document before executing the price lookup.

10. One conversation, several kinds of computation

The diagram shows a single conversation moving through several representations:

This minimal voice agent helped me understand how each transformation makes the request flow through the pipeline. Whisper transcribes audio, MiniLM vectorizes the text, SmolLM2 predicts the tool call, Rust executes it, and Piper sends the reply back to the user. It’s fun to observe and debug this voice agent and it’s moving parts.

Final Thoughts

I have been struggling with the deluge of information from the past year, and it is hard to keep up. In such cases, whenever I’m drowning in information, I peel the layer and go back to the first principles to reorient myself. An afternoon of vibe engineering, and imagining voice agents from the first principles helped me understand how all these stages and models come together to provide a seamless experience.

Just like what Jacobi and Munger said.

Footnotes #

Stanford Introduction to Information Retrieval: Dot products explains cosine similarity between vector representations.↩ 2. Whisper audio preprocessing uses mono audio sampled at 16 kHz.↩ 3. OpenAI Whisper README describes the model’s training and speech recognition tasks.↩ 4. Piper documentation describes its phonemization and ONNX voice model↩

── more in #ai-agents 4 stories · sorted by recency
── more on @voxlocal 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/voxlocal-a-minimal-v…] indexed:0 read:15min 2026-10-10 · —