cd /news/ai-tools/agent-voice-kit-local-talk-with-your… · home topics ai-tools article
[ARTICLE · art-119223] src=github.com ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Agent Voice Kit: Local talk with your Agents (whisper+supertonic)

Francisco Carlos Serra released Agent Voice Kit, an open-source tool that lets users talk to AI agents like Claude Code, pi, Codex, and Gemini CLI locally, with dictation and optional text-to-speech, running entirely on the user's computer. The kit requires Python 3.11+, ffmpeg, and a GPU for optimal performance, and downloads Whisper large-v3-turbo and Supertonic 3 models from Hugging Face (about 2 GB).

read2 min views2 publishedSep 2, 2026
Agent Voice Kit: Local talk with your Agents (whisper+supertonic)
Image: Michielbdejong (auto-discovered)

English · Español

you speak                  the agent answers              you hear it (optional)
press F1, talk, press F1   the text lands in the window   the reply is read aloud

Press a key, say what you want, press it again. Your words are typed wherever you were writing, in any app. If you want, the agent reads its answer back to you.

simplescreenrecorder-2026-09-02_12.44.15_audio_v3_1.5x.mp4 #

It all runs on your own computer. Nothing is sent anywhere, and there is no waiting for a server: the text is there the moment you stop talking.

Paste this into your agent (Claude Code, pi, Codex, Gemini CLI):

Install https://github.com/franciscocarloserra/agent-voice-kit/ following its AGENTS.md

What you will have to do by hand. The agent cannot click these for you:

  • Mac: type your password for brew

, assign the two hotkeys in System Settings, accept the Microphone and Accessibility dialogs. - Windows: approve the ffmpeg and AutoHotkey installs when winget asks.

  • Linux: type your sudo

password for the packages, assign the two hotkeys in your desktop's shortcut settings.

key what it does
F1 starts recording; pressing F1 again types your words at the cursor, in any app
Meta+F1 cancels the recording or silences the agent
/tts
turns read-aloud on or off, in any harness. Persists until you toggle it again

python3 voice.py serve

runs the servers in a terminal, install-service

runs them at login, doctor

reports what is missing, and tts speed 1.5

sets the reading speed.

You want a GPU. NVIDIA on Linux or Windows, or any Apple Silicon Mac. On CPU only it works but takes a few seconds per sentence.

Models are downloaded, not bundled. Setup pulls Whisper large-v3-turbo and Supertonic 3 from Hugging Face, about 2 GB.

A few things get installed on your system. Python 3.11+ and ffmpeg everywhere, AutoHotkey on Windows, xdotool on Linux, current NVIDIA drivers if you have that GPU.

Text to speech works out of the box, and can get faster. The default engine needs no build. If you want replies to start instantly, ask the agent to build audio.cpp (about 15 minutes, 5 to 8 times faster).

macOS will ask for permissions. Microphone and Accessibility, once. Without them dictation records silence or does not type.

Read-aloud is off until you turn it on. /tts

toggles it for every harness and it stays that way. Dictation needs nothing: it types into whatever window has focus.

Everything is customizable. Models, voice, reading speed, hotkeys, ports. Just ask your agent to change it; it all lives in config.json

.

── more in #ai-tools 4 stories · sorted by recency
── more on @francisco carlos serra 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agent-voice-kit-loca…] indexed:0 read:2min 2026-09-02 ·