cd /news/ai-agents/show-hn-talktome-voice-calls-with-yo… · home › topics › ai-agents › article
[ARTICLE · art-141389] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Show HN: Talktome – Voice calls with your AI agent

Developer rohanprichard released TalkToMe, a macOS menu-bar app that lets users hold voice calls with an existing coding agent session, with the agent placing the call via the `talktome call` command using its session ID. TalkToMe supplies no language model of its own: it records the microphone, converts speech to text, sends the text into the agent's session, and speaks the agent's reply, supporting Codex, Claude Code, Hermes Agent, OpenClaw, or any host that can run shell commands. The app requires macOS on Apple silicon (arm64), Node.js 22 or later, uv (which installs Python 3.11 to 3.13), Xcode command-line tools, and an optional ElevenLabs API key for cloud speech or local speech otherwise.

read8 min views1 publishedSep 29, 2026
Show HN: Talktome – Voice calls with your AI agent
Image: Michielbdejong (auto-discovered)

TalkToMe is a macOS menu-bar app for voice calls with a coding agent. The agent rings you from the session that it already runs. You answer, and you talk.

The agent keeps its model, tools, files, and history. TalkToMe does not supply a language model. The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply.

  1. You ask the agent in its session to call you, for example "Call me with TalkToMe."
  2. The agent runs talktome call with its session ID.
  3. TalkToMe shows a ring on the screen. You select Answer .
  4. The agent speaks a short greeting.
  5. You speak. The agent receives your words as a message in its session and replies.
  6. You select End , or the agent runstalktome end .

The app has no window that starts a conversation. After setup, TalkToMe stays in the menu bar and waits for a call. The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call.

  • macOS. The app received tests only on macOS. The disk image is for Apple silicon (arm64).
  • Node.js 22 or later, for a build from source.
  • uv . uv installs Python 3.11 to 3.13 when necessary.
  • Xcode command-line tools, to build the disk image. The build compiles a small Swift helper.
  • One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands.
git clone https://github.com/rohanprichard/talktome.git
cd talktome
npm ci
uv sync --frozen
npm start

npm start runs uv sync --frozen if the Python environment is missing. Then it starts Electron.

The first start opens a setup window with six steps:

  1. Welcome. The window explains the call flow.
  2. Microphone. SelectAllow microphone . macOS asks for permission.
  3. ElevenLabs. Enter an ElevenLabs API key, or selectLater to use local speech.
  4. Agent connection. SelectInstall . This installs the agent skill and thetalktome command.
  5. Glow color. Select the color that the call surface shows during a live call.
  6. All set. Ask your agent to call.

You can skip a step with Later. Settings contains the same options. To run setup again, quit the app and run npm run reset. This command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token.

Install writes the skill file to ~/.codex/skills/talktome/SKILL.md. It also writes the skill to ~/.hermes/skills/ and ~/.openclaw/skills/ if those directories exist. It installs the talktome command in ~/.local/bin, /opt/homebrew/bin, or /usr/local/bin. The skill tells the agent which commands to run. See the skill.

For Claude Code, or for another host, copy the skill into the host's skill directory:

mkdir -p ~/.claude/skills/talktome
talktome skill > ~/.claude/skills/talktome/SKILL.md
Host Command How replies reach the call
Codex talktome call --agent codex --thread "$CODEX_THREAD_ID" TalkToMe reads the session's public replies automatically.
Claude Code talktome call --agent claude --thread SESSION_ID The agent runs talktome listen andtalktome reply .
Hermes Agent chat talktome call --agent hermes --connection cooperative --thread ID The agent runs talktome listen andtalktome reply .
OpenClaw chat talktome call --agent openclaw --connection cooperative --thread ID The agent runs talktome listen andtalktome reply .
Other hosts talktome call --agent generic --thread ID The agent runs talktome listen andtalktome reply .

The listen and reply commands are the "cooperative" connection. They work with any host that can run shell commands on the Mac that runs TalkToMe. The commands exchange private files with the app. Thus, they work when a sandbox blocks local network access.

Hermes and OpenClaw also have experimental adapters for an API session or a Gateway session. Run talktome providers to see the connection methods that are ready. Agent support gives the setup, the limits, and the interruption behavior.

Agent protocol gives the commands, the ring flow, and the local HTTP interface.

The remote bridge lets an agent on another server ring the laptop. Only text and call events cross the bridge. Microphone audio stays on the laptop. The setup uses SSH:

  1. Install talktome on the server: uv tool install git+https://github.com/rohanprichard/talktome .
  2. On the laptop, run talktome remote-connect user@server --install-service .
  3. Restart TalkToMe.

The app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt. This feature is experimental. See remote bridge.

The microphone stays active while the agent speaks. If you speak during a reply, the reply s. If the app hears words, it clears the old reply and starts a new turn. If it hears no words, the reply continues. The agent receives a short report of how much of its last reply played. Codex and cooperative hosts keep control of their work. An interruption does not stop a tool that already started.

TalkToMe uses Smart Turn, a small local model, to decide when you finished speaking. Settings can select a fixed instead. See Smart Turn. Call latency and call timing describe the delays in a call.

Settings also sets the position of the call surface: Bottom or Top center.

Function Local option ElevenLabs option
Speech recognition Whisper Small (484 MB) or Whisper Base English (145 MB) Scribe v2
Agent voice System voice or Kokoro Flash v2.5

Whisper is the default for recognition. The system voice is the default voice. Download a Whisper model in Settings before you use local recognition. Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use. Local models run offline after the download.

The ElevenLabs key needs access to the voice list and to each selected speech service. Select Remember key to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes. The app never returns the key to the interface or writes it to its settings file. Remove key removes the key from the app session and from the keychain. See speech providers for the exact interfaces.

The server listens only on 127.0.0.1:8765. A generated local token protects its interface. The desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface. The app keeps its token, settings, and models in ~/Library/Application Support/talktome. The app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them.

The app connects to the network for these purposes only:

  • Hugging Face, to download Whisper and the Smart Turn model.
  • GitHub, to download the Kokoro model and voice file.
  • ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply.
  • A Hermes or OpenClaw host, or a relay, only if you configure one.

The agent receives the text of what you say. The agent's provider and tools have their own data rules. The cooperative command files contain conversation text. Settings for Hermes and OpenClaw are in agent-hosts.json in the data directory. This file holds a plaintext token with mode 0600.

Environment settings:

Name Purpose
TALKTOME_DATA_DIR Change the local data directory
TALKTOME_PORT Change the desktop server port
TALKTOME_URL Set the server address for external clients
TALKTOME_TOKEN Supply an existing shared token
TALKTOME_RELOAD Restart the server when Python files change. Development only.
TALKTOME_FLOATING_CALL Set to 0 to turn off the call window. The call then has no controls on screen.
TALKTOME_ALLOW_REMOTE_AGENTS Set to 1 to permit a Hermes host that is not on this Mac. The URL must use HTTPS.
npm run build:app

This command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder. The result is dist/app/TalkToMe-<version>-arm64.dmg. The script mounts the disk image after the build and examines its contents. To reuse the last frozen server when only the desktop code changed, run npm run build:dmg.

To install the app, open the disk image and drag TalkToMe to Applications. The build is not signed. It opens on the Mac that built it. Other Macs block it, because notarization needs a paid Apple Developer ID. The bundle does not include the speech models. The app downloads them at first use.

uv sync --frozen        # install the Python environment
npm ci                  # install the Node packages
npm start               # start the app from source
uv run pytest -q        # Python tests
uv run ruff check       # Python lint
node --test tests/      # JavaScript tests

npm test runs the Python tests and the JavaScript tests together. These commands start the real app for end-to-end checks:

Command What it examines
npm run test:call The call window: position, stacking, and growth of the transcript
npm run test:attach-call A full call against a real Codex session. It sends a few short model requests.
npm run test:stream Time to first audio for ElevenLabs. It spends credits on two short replies.

The call test turns off the Chromium sandbox. A normal start keeps the sandbox on.

The development notes hold plans, research, and a work log. They can be out of date.

To report a security problem, read SECURITY.md.

Read CONTRIBUTING.md before you open a pull request.

MIT. See LICENSE. TalkToMe uses Electron, FastAPI, faster-whisper, kokoro-onnx, Pipecat Smart Turn, and optional ElevenLabs services. It does not contain copied SpeakType or AgentCall code. NOTICE lists the third-party references.

── more in #ai-agents 4 stories · sorted by recency
── more on @talktome 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-talktome-voi…] indexed:0 read:8min 2026-09-29 · —