{"slug": "show-hn-talktome-voice-calls-with-your-ai-agent", "title": "Show HN: Talktome – Voice calls with your AI agent", "summary": "Developer rohanprichard released TalkToMe, a macOS menu-bar app that lets users hold voice calls with an existing coding agent session, with the agent placing the call via the `talktome call` command using its session ID. TalkToMe supplies no language model of its own: it records the microphone, converts speech to text, sends the text into the agent's session, and speaks the agent's reply, supporting Codex, Claude Code, Hermes Agent, OpenClaw, or any host that can run shell commands. The app requires macOS on Apple silicon (arm64), Node.js 22 or later, uv (which installs Python 3.11 to 3.13), Xcode command-line tools, and an optional ElevenLabs API key for cloud speech or local speech otherwise.", "body_md": "TalkToMe is a macOS menu-bar app for voice calls with a coding agent. The agent rings you from the session that it already runs. You answer, and you talk.\n\nThe agent keeps its model, tools, files, and history. TalkToMe does not supply a language model. The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply.\n\n1. You ask the agent in its session to call you, for example \"Call me with TalkToMe.\"\n2. The agent runs `talktome call` with its session ID.\n3. TalkToMe shows a ring on the screen. You select **Answer** .\n4. The agent speaks a short greeting.\n5. You speak. The agent receives your words as a message in its session and replies.\n6. You select **End** , or the agent runs`talktome end` .\n\nThe app has no window that starts a conversation. After setup, TalkToMe stays in the menu bar and waits for a call. The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call.\n\n- macOS. The app received tests only on macOS. The disk image is for Apple silicon (arm64).\n- Node.js 22 or later, for a build from source.\n- [uv](https://docs.astral.sh/uv/getting-started/installation/) . uv installs Python 3.11 to 3.13 when necessary.\n- Xcode command-line tools, to build the disk image. The build compiles a small Swift helper.\n- One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands.\n\n```\ngit clone https://github.com/rohanprichard/talktome.git\ncd talktome\nnpm ci\nuv sync --frozen\nnpm start\n```\n\n`npm start` runs `uv sync --frozen` if the Python environment is missing. Then it starts Electron.\n\nThe first start opens a setup window with six steps:\n\n1. **Welcome.** The window explains the call flow.\n2. **Microphone.** Select**Allow microphone** . macOS asks for permission.\n3. **ElevenLabs.** Enter an ElevenLabs API key, or select**Later** to use local speech.\n4. **Agent connection.** Select**Install** . This installs the agent skill and the`talktome` command.\n5. **Glow color.** Select the color that the call surface shows during a live call.\n6. **All set.** Ask your agent to call.\n\nYou can skip a step with **Later**. Settings contains the same options.\nTo run setup again, quit the app and run `npm run reset`.\nThis command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token.\n\n**Install** writes the skill file to `~/.codex/skills/talktome/SKILL.md`.\nIt also writes the skill to `~/.hermes/skills/` and `~/.openclaw/skills/` if those directories exist.\nIt installs the `talktome` command in `~/.local/bin`, `/opt/homebrew/bin`, or `/usr/local/bin`.\nThe skill tells the agent which commands to run. See [the skill](https://github.com/rohanprichard/talktome/blob/main/skills/talktome/SKILL.md).\n\nFor Claude Code, or for another host, copy the skill into the host's skill directory:\n\n```\nmkdir -p ~/.claude/skills/talktome\ntalktome skill > ~/.claude/skills/talktome/SKILL.md\n```\n\n| Host | Command | How replies reach the call | \n|---|---|---|\n| Codex | `talktome call --agent codex --thread \"$CODEX_THREAD_ID\"` | TalkToMe reads the session's public replies automatically. | \n| Claude Code | `talktome call --agent claude --thread SESSION_ID` | The agent runs `talktome listen` and`talktome reply` . | \n| Hermes Agent chat | `talktome call --agent hermes --connection cooperative --thread ID` | The agent runs `talktome listen` and`talktome reply` . | \n| OpenClaw chat | `talktome call --agent openclaw --connection cooperative --thread ID` | The agent runs `talktome listen` and`talktome reply` . | \n| Other hosts | `talktome call --agent generic --thread ID` | The agent runs `talktome listen` and`talktome reply` . | \n\nThe `listen` and `reply` commands are the \"cooperative\" connection.\nThey work with any host that can run shell commands on the Mac that runs TalkToMe.\nThe commands exchange private files with the app. Thus, they work when a sandbox blocks local network access.\n\nHermes and OpenClaw also have experimental adapters for an API session or a Gateway session.\nRun `talktome providers` to see the connection methods that are ready.\n[Agent support](https://github.com/rohanprichard/talktome/blob/main/docs/AGENT_SUPPORT.md) gives the setup, the limits, and the interruption behavior.\n\n[Agent protocol](https://github.com/rohanprichard/talktome/blob/main/docs/AGENT_API.md) gives the commands, the ring flow, and the local HTTP interface.\n\nThe remote bridge lets an agent on another server ring the laptop. Only text and call events cross the bridge. Microphone audio stays on the laptop. The setup uses SSH:\n\n1. Install talktome on the server: `uv tool install git+https://github.com/rohanprichard/talktome` .\n2. On the laptop, run `talktome remote-connect user@server --install-service` .\n3. Restart TalkToMe.\n\nThe app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt.\nThis feature is experimental. See [remote bridge](https://github.com/rohanprichard/talktome/blob/main/docs/REMOTE_BRIDGE.md).\n\nThe microphone stays active while the agent speaks. If you speak during a reply, the reply pauses. If the app hears words, it clears the old reply and starts a new turn. If it hears no words, the reply continues. The agent receives a short report of how much of its last reply played. Codex and cooperative hosts keep control of their work. An interruption does not stop a tool that already started.\n\nTalkToMe uses Smart Turn, a small local model, to decide when you finished speaking.\nSettings can select a fixed pause instead. See [Smart Turn](https://github.com/rohanprichard/talktome/blob/main/docs/SMART_TURN.md).\n[Call latency](https://github.com/rohanprichard/talktome/blob/main/docs/LATENCY.md) and [call timing](https://github.com/rohanprichard/talktome/blob/main/docs/CALL_TIMING.md) describe the delays in a call.\n\nSettings also sets the position of the call surface: **Bottom** or **Top center**.\n\n| Function | Local option | ElevenLabs option | \n|---|---|---|\n| Speech recognition | Whisper Small (484 MB) or Whisper Base English (145 MB) | Scribe v2 | \n| Agent voice | System voice or Kokoro | Flash v2.5 | \n\nWhisper is the default for recognition. The system voice is the default voice. Download a Whisper model in Settings before you use local recognition. Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use. Local models run offline after the download.\n\nThe ElevenLabs key needs access to the voice list and to each selected speech service.\nSelect **Remember key** to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes.\nThe app never returns the key to the interface or writes it to its settings file.\n**Remove key** removes the key from the app session and from the keychain.\nSee [speech providers](https://github.com/rohanprichard/talktome/blob/main/docs/SPEECH_PROVIDERS.md) for the exact interfaces.\n\nThe server listens only on `127.0.0.1:8765`. A generated local token protects its interface.\nThe desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface.\nThe app keeps its token, settings, and models in `~/Library/Application Support/talktome`.\nThe app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them.\n\nThe app connects to the network for these purposes only:\n\n- Hugging Face, to download Whisper and the Smart Turn model.\n- GitHub, to download the Kokoro model and voice file.\n- ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply.\n- A Hermes or OpenClaw host, or a relay, only if you configure one.\n\nThe agent receives the text of what you say. The agent's provider and tools have their own data rules.\nThe cooperative command files contain conversation text.\nSettings for Hermes and OpenClaw are in `agent-hosts.json` in the data directory. This file holds a plaintext token with mode `0600`.\n\nEnvironment settings:\n\n| Name | Purpose | \n|---|---|\n| `TALKTOME_DATA_DIR` | Change the local data directory | \n| `TALKTOME_PORT` | Change the desktop server port | \n| `TALKTOME_URL` | Set the server address for external clients | \n| `TALKTOME_TOKEN` | Supply an existing shared token | \n| `TALKTOME_RELOAD` | Restart the server when Python files change. Development only. | \n| `TALKTOME_FLOATING_CALL` | Set to `0` to turn off the call window. The call then has no controls on screen. | \n| `TALKTOME_ALLOW_REMOTE_AGENTS` | Set to `1` to permit a Hermes host that is not on this Mac. The URL must use HTTPS. | \n\n```\nnpm run build:app\n```\n\nThis command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder.\nThe result is `dist/app/TalkToMe-<version>-arm64.dmg`.\nThe script mounts the disk image after the build and examines its contents.\nTo reuse the last frozen server when only the desktop code changed, run `npm run build:dmg`.\n\nTo install the app, open the disk image and drag TalkToMe to Applications. The build is not signed. It opens on the Mac that built it. Other Macs block it, because notarization needs a paid Apple Developer ID. The bundle does not include the speech models. The app downloads them at first use.\n\n```\nuv sync --frozen        # install the Python environment\nnpm ci                  # install the Node packages\nnpm start               # start the app from source\nuv run pytest -q        # Python tests\nuv run ruff check       # Python lint\nnode --test tests/      # JavaScript tests\n```\n\n`npm test` runs the Python tests and the JavaScript tests together.\nThese commands start the real app for end-to-end checks:\n\n| Command | What it examines | \n|---|---|\n| `npm run test:call` | The call window: position, stacking, and growth of the transcript | \n| `npm run test:attach-call` | A full call against a real Codex session. It sends a few short model requests. | \n| `npm run test:stream` | Time to first audio for ElevenLabs. It spends credits on two short replies. | \n\nThe call test turns off the Chromium sandbox. A normal start keeps the sandbox on.\n\nThe [development notes](https://github.com/rohanprichard/talktome/blob/main/docs/notes/README.md) hold plans, research, and a work log. They can be out of date.\n\nTo report a security problem, read [SECURITY.md](https://github.com/rohanprichard/talktome/blob/main/SECURITY.md).\n\nRead [CONTRIBUTING.md](https://github.com/rohanprichard/talktome/blob/main/CONTRIBUTING.md) before you open a pull request.\n\nMIT. See [LICENSE](https://github.com/rohanprichard/talktome/blob/main/LICENSE).\nTalkToMe uses Electron, FastAPI, faster-whisper, kokoro-onnx, Pipecat Smart Turn, and optional ElevenLabs services.\nIt does not contain copied SpeakType or AgentCall code. [NOTICE](https://github.com/rohanprichard/talktome/blob/main/NOTICE) lists the third-party references.", "url": "https://wpnews.pro/news/show-hn-talktome-voice-calls-with-your-ai-agent", "canonical_source": "https://github.com/rohanprichard/talktome", "published_at": "2026-09-29 02:01:08+00:00", "updated_at": "2026-09-29 02:17:40.359551+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "ai-products", "developer-tools"], "entities": ["TalkToMe", "rohanprichard", "Codex", "Claude Code", "Hermes Agent", "OpenClaw", "ElevenLabs", "macOS"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-talktome-voice-calls-with-your-ai-agent", "markdown": "https://wpnews.pro/news/show-hn-talktome-voice-calls-with-your-ai-agent.md", "text": "https://wpnews.pro/news/show-hn-talktome-voice-calls-with-your-ai-agent.txt", "jsonld": "https://wpnews.pro/news/show-hn-talktome-voice-calls-with-your-ai-agent.jsonld"}}