Show HN: Talktome – Voice calls with your AI agent Developer rohanprichard released TalkToMe, a macOS menu-bar app that lets users hold voice calls with an existing coding agent session, with the agent placing the call via the `talktome call` command using its session ID. TalkToMe supplies no language model of its own: it records the microphone, converts speech to text, sends the text into the agent's session, and speaks the agent's reply, supporting Codex, Claude Code, Hermes Agent, OpenClaw, or any host that can run shell commands. The app requires macOS on Apple silicon (arm64), Node.js 22 or later, uv (which installs Python 3.11 to 3.13), Xcode command-line tools, and an optional ElevenLabs API key for cloud speech or local speech otherwise. TalkToMe is a macOS menu-bar app for voice calls with a coding agent. The agent rings you from the session that it already runs. You answer, and you talk. The agent keeps its model, tools, files, and history. TalkToMe does not supply a language model. The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply. 1. You ask the agent in its session to call you, for example "Call me with TalkToMe." 2. The agent runs talktome call with its session ID. 3. TalkToMe shows a ring on the screen. You select Answer . 4. The agent speaks a short greeting. 5. You speak. The agent receives your words as a message in its session and replies. 6. You select End , or the agent runs talktome end . The app has no window that starts a conversation. After setup, TalkToMe stays in the menu bar and waits for a call. The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call. - macOS. The app received tests only on macOS. The disk image is for Apple silicon arm64 . - Node.js 22 or later, for a build from source. - uv https://docs.astral.sh/uv/getting-started/installation/ . uv installs Python 3.11 to 3.13 when necessary. - Xcode command-line tools, to build the disk image. The build compiles a small Swift helper. - One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands. git clone https://github.com/rohanprichard/talktome.git cd talktome npm ci uv sync --frozen npm start npm start runs uv sync --frozen if the Python environment is missing. Then it starts Electron. The first start opens a setup window with six steps: 1. Welcome. The window explains the call flow. 2. Microphone. Select Allow microphone . macOS asks for permission. 3. ElevenLabs. Enter an ElevenLabs API key, or select Later to use local speech. 4. Agent connection. Select Install . This installs the agent skill and the talktome command. 5. Glow color. Select the color that the call surface shows during a live call. 6. All set. Ask your agent to call. You can skip a step with Later . Settings contains the same options. To run setup again, quit the app and run npm run reset . This command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token. Install writes the skill file to ~/.codex/skills/talktome/SKILL.md . It also writes the skill to ~/.hermes/skills/ and ~/.openclaw/skills/ if those directories exist. It installs the talktome command in ~/.local/bin , /opt/homebrew/bin , or /usr/local/bin . The skill tells the agent which commands to run. See the skill https://github.com/rohanprichard/talktome/blob/main/skills/talktome/SKILL.md . For Claude Code, or for another host, copy the skill into the host's skill directory: mkdir -p ~/.claude/skills/talktome talktome skill ~/.claude/skills/talktome/SKILL.md | Host | Command | How replies reach the call | |---|---|---| | Codex | talktome call --agent codex --thread "$CODEX THREAD ID" | TalkToMe reads the session's public replies automatically. | | Claude Code | talktome call --agent claude --thread SESSION ID | The agent runs talktome listen and talktome reply . | | Hermes Agent chat | talktome call --agent hermes --connection cooperative --thread ID | The agent runs talktome listen and talktome reply . | | OpenClaw chat | talktome call --agent openclaw --connection cooperative --thread ID | The agent runs talktome listen and talktome reply . | | Other hosts | talktome call --agent generic --thread ID | The agent runs talktome listen and talktome reply . | The listen and reply commands are the "cooperative" connection. They work with any host that can run shell commands on the Mac that runs TalkToMe. The commands exchange private files with the app. Thus, they work when a sandbox blocks local network access. Hermes and OpenClaw also have experimental adapters for an API session or a Gateway session. Run talktome providers to see the connection methods that are ready. Agent support https://github.com/rohanprichard/talktome/blob/main/docs/AGENT SUPPORT.md gives the setup, the limits, and the interruption behavior. Agent protocol https://github.com/rohanprichard/talktome/blob/main/docs/AGENT API.md gives the commands, the ring flow, and the local HTTP interface. The remote bridge lets an agent on another server ring the laptop. Only text and call events cross the bridge. Microphone audio stays on the laptop. The setup uses SSH: 1. Install talktome on the server: uv tool install git+https://github.com/rohanprichard/talktome . 2. On the laptop, run talktome remote-connect user@server --install-service . 3. Restart TalkToMe. The app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt. This feature is experimental. See remote bridge https://github.com/rohanprichard/talktome/blob/main/docs/REMOTE BRIDGE.md . The microphone stays active while the agent speaks. If you speak during a reply, the reply pauses. If the app hears words, it clears the old reply and starts a new turn. If it hears no words, the reply continues. The agent receives a short report of how much of its last reply played. Codex and cooperative hosts keep control of their work. An interruption does not stop a tool that already started. TalkToMe uses Smart Turn, a small local model, to decide when you finished speaking. Settings can select a fixed pause instead. See Smart Turn https://github.com/rohanprichard/talktome/blob/main/docs/SMART TURN.md . Call latency https://github.com/rohanprichard/talktome/blob/main/docs/LATENCY.md and call timing https://github.com/rohanprichard/talktome/blob/main/docs/CALL TIMING.md describe the delays in a call. Settings also sets the position of the call surface: Bottom or Top center . | Function | Local option | ElevenLabs option | |---|---|---| | Speech recognition | Whisper Small 484 MB or Whisper Base English 145 MB | Scribe v2 | | Agent voice | System voice or Kokoro | Flash v2.5 | Whisper is the default for recognition. The system voice is the default voice. Download a Whisper model in Settings before you use local recognition. Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use. Local models run offline after the download. The ElevenLabs key needs access to the voice list and to each selected speech service. Select Remember key to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes. The app never returns the key to the interface or writes it to its settings file. Remove key removes the key from the app session and from the keychain. See speech providers https://github.com/rohanprichard/talktome/blob/main/docs/SPEECH PROVIDERS.md for the exact interfaces. The server listens only on 127.0.0.1:8765 . A generated local token protects its interface. The desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface. The app keeps its token, settings, and models in ~/Library/Application Support/talktome . The app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them. The app connects to the network for these purposes only: - Hugging Face, to download Whisper and the Smart Turn model. - GitHub, to download the Kokoro model and voice file. - ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply. - A Hermes or OpenClaw host, or a relay, only if you configure one. The agent receives the text of what you say. The agent's provider and tools have their own data rules. The cooperative command files contain conversation text. Settings for Hermes and OpenClaw are in agent-hosts.json in the data directory. This file holds a plaintext token with mode 0600 . Environment settings: | Name | Purpose | |---|---| | TALKTOME DATA DIR | Change the local data directory | | TALKTOME PORT | Change the desktop server port | | TALKTOME URL | Set the server address for external clients | | TALKTOME TOKEN | Supply an existing shared token | | TALKTOME RELOAD | Restart the server when Python files change. Development only. | | TALKTOME FLOATING CALL | Set to 0 to turn off the call window. The call then has no controls on screen. | | TALKTOME ALLOW REMOTE AGENTS | Set to 1 to permit a Hermes host that is not on this Mac. The URL must use HTTPS. | npm run build:app This command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder. The result is dist/app/TalkToMe-