{"slug": "build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api", "title": "Build a voice agent with Whisper, Kokoro, and an OpenAI-compatible API", "summary": "A developer walkthrough shows how to assemble a full voice agent — speech-to-text, LLM, and text-to-speech — using three OpenAI-compatible API calls on EcoHash with a single key and base URL. The pipeline pairs whisper-large-v3-turbo for transcription, llama-3.1-8b-instruct for replies, and kokoro-82m for speech output at roughly $0.006 per minute of generated speech, with no local GPU required.", "body_md": "A voice agent turns speech into speech: it transcribes what the user says, sends the text to a language model, and speaks the reply back. On EcoHash you build all three stages through one OpenAI-compatible API and one key, so there are no three vendors and no three billing accounts to stitch together. Whisper (`whisper-large-v3-turbo`) does speech to text, a chat model such as `llama-3.1-8b-instruct` writes the reply, and Kokoro (`kokoro-82m`) turns it into audio. None of it needs a GPU of your own, since the models are served for you. This post walks through the pipeline, the three API calls, and when an RTX Pro 6000 workspace is worth it for heavier voice work.\n\n**Three API calls, one key.** Whisper hears, a chat model thinks, Kokoro speaks. About $0.006 per minute of generated speech.\n\nThe loop has three stages, and each maps to one EcoHash endpoint:\n\nThen you repeat for the next turn. The whole loop uses one key and one base URL, so the three stages share auth, billing, and SDK.\n\n`whisper-large-v3-turbo`\n`qwen3-asr-1-7b`\n`llama-3.1-8b-instruct`\n`kokoro-82m`\n`qwen3-tts`\nUse a small or mid-size chat model for the LLM stage so replies come back fast. A coding-specialized model is the wrong pick for general conversation.\n\nHere is what the speech models measure on a single RTX Pro 6000, end-to-end, in July 2026. Speech to text:\n\nText to speech:\n\nOn the HF Open ASR Leaderboard, the models EcoHash serves land among the accurate ones (purple = served on EcoHash):\n\nFull data and method: [ecohash-benchmarks](https://github.com/ecohash-ai/ecohash-benchmarks).\n\nThe API is OpenAI-compatible, so the OpenAI SDK works once you set the base URL and key. One client covers all three stages.\n\n``` python\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"https://api.ecohash.com/v1\",\n    api_key=\"eco_...\",  # create a key at console.ecohash.com\n)\n\n# 1. Speech to text\nwith open(\"input.wav\", \"rb\") as f:\n    transcript = client.audio.transcriptions.create(\n        model=\"whisper-large-v3-turbo\",\n        file=f,\n    ).text\n\n# 2. LLM reply\nreply = client.chat.completions.create(\n    model=\"llama-3.1-8b-instruct\",\n    messages=[\n        {\"role\": \"system\", \"content\": \"You are a concise voice assistant. Keep replies short.\"},\n        {\"role\": \"user\", \"content\": transcript},\n    ],\n).choices[0].message.content\n\n# 3. Text to speech\nwith client.audio.speech.with_streaming_response.create(\n    model=\"kokoro-82m\",\n    voice=\"af_heart\",\n    input=reply,\n    response_format=\"wav\",\n) as speech:\n    speech.stream_to_file(\"reply.wav\")\n\nprint(\"You said:\", transcript)\nprint(\"Assistant:\", reply)\n```\n\nThat is a full single-turn voice agent. For multi-turn, keep the message history and run the loop again for each new audio input.\n\nYou do not need one. All three models are served through the API, so you can build and run a voice agent with no hardware to provision, and Kokoro in particular is small enough that it never needs an RTX Pro 6000.\n\nA workspace helps later, not at the start: when you want to co-locate a heavier speech-to-text, LLM, and text-to-speech pipeline on one card, push higher throughput, or run your own serving stack. That is a scaling choice. See [/gpu-compute](https://ecohash.com/gpu-compute).\n\n**What is a voice agent?** A system that listens, thinks, and speaks: speech to text, then an LLM, then text to speech. The user talks, and the agent replies in audio.\n\n**Which models do I need?** Three: a speech-to-text model like `whisper-large-v3-turbo`, a chat model like `llama-3.1-8b-instruct`, and a text-to-speech model like `kokoro-82m`.\n\n**Can I do all of it with one API key?** Yes. On EcoHash the transcription, chat, and speech endpoints share one base URL and one key, so there is no three-vendor integration.\n\n**Do I need an RTX Pro 6000 to run Kokoro or Whisper?** No. Both run through the API. A workspace is optional, for co-locating a heavier pipeline or higher throughput.\n\n**Is it fast enough for real-time conversation?** That depends on your network, model choices, streaming, and pipeline. This post does not publish latency numbers; measure your own setup before relying on a target.\n\n**Which voices does Kokoro support?** Several; the example uses `af_heart`. See [/models/kokoro-82m](https://ecohash.com/models/kokoro-82m).\n\n**How much does a voice agent cost?** The LLM stage is billed per token, and the audio stages by what they process. See [/pricing](https://ecohash.com/pricing) for current rates.\n\n**Which chat model should I use for the LLM stage?** A small or mid-size general model keeps replies quick. See [Best models to run on RTX Pro 6000 96GB](https://ecohash.com/blog/best-models-for-rtx-pro-6000).", "url": "https://wpnews.pro/news/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api", "canonical_source": "https://dev.to/ecohash/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api-1ee", "published_at": "2026-10-10 07:03:48+00:00", "updated_at": "2026-10-10 07:10:28.773807+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "natural-language-processing", "ai-infrastructure"], "entities": ["EcoHash", "Whisper", "Kokoro", "OpenAI", "llama-3.1-8b-instruct", "whisper-large-v3-turbo", "kokoro-82m", "RTX Pro 6000"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api", "markdown": "https://wpnews.pro/news/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api.md", "text": "https://wpnews.pro/news/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api.txt", "jsonld": "https://wpnews.pro/news/build-a-voice-agent-with-whisper-kokoro-and-an-openai-compatible-api.jsonld"}}