{"slug": "gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent", "title": "Gemini 3.8 Live is finally in the API and the pricing is actually decent", "summary": "Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking in its API and Google AI Studio, adding native speech-to-speech capability for voice-first agents. Gemini 3.8 Live Extended Thinking currently ranks #1 on the Artificial Analysis Speech-to-Speech leaderboard, while Gemini 3.5 Transcribe supports over 85 languages with a 2.6% Word Error Rate for non-streaming and 4.0% for streaming. Live API pricing is $0.005/min for audio input and $0.018/min for audio output, with integrations from LiveKit, Vercel, LangChain, Agora, Pipecat, Fishjam, and Vision Agents.", "body_md": "# Gemini 3.8 Live is finally in the API and the pricing is actually decent\n\nMy team has been tasked with moving our voice-first features away from clunky cascaded architectures, and the release of [Gemini](https://promptcube3.com/en/tags/gemini/) 3.8 Live and 3.8 Live Extended Thinking in the API and Google AI Studio makes that a lot easier. The big win here is the native speech-to-speech capability, which means the model can actually handle tasks without killing the flow of the conversation.\n\n## Which model fits which use case?\n\nDepending on how much \"brain power\" the agent needs, there are two paths for the live experience. Gemini 3.8 Live is the standard for maintaining dialogue while performing tasks. If the request is actually complex, Gemini 3.8 Live Extended Thinking is the move—it's currently ranking #1 on the Artificial Analysis Speech-to-Speech leaderboard.\n\nFor those of us who just need raw text from audio, Gemini 3.5 Transcribe is the dedicated tool. It supports over 85 languages and the accuracy is impressive, hitting a 2.6% Word Error Rate (WER) for non-streaming and 4.0% for streaming.\n\n## What actually changes for the dev workflow?\n\nIntegrating these into a product changes a few things about how we handle agent logic. A few specific features that stand out for our current rollout:\n\n- **Asynchronous function calling:** This is huge. The agent can trigger API or tool calls in the background while it keeps streaming audio to the user. No more awkward silence while the bot \"thinks\" or fetches data.\n- **Visual context:** The models can ground the dialogue in live visual inputs, so the agent sees what the user sees.\n- **Alphanumeric precision:** It's actually reliable at parsing things like claim numbers or confirmation codes, which usually get mangled in voice apps.\n- **Incremental content updates:** It can merge real-time audio with structured data for context-aware responses.\n- **Language support:** It covers 97+ languages with consistent accents.\n\nIf you use the Extended Thinking version, you can use configurable thinking to handle multi-step reasoning in the background while the model narrates its progress to the user.\n\n## The cost and infrastructure side\n\nThe pricing for the Live API is straightforward:\n\n- **Audio Input:** $0.005/min\n- **Audio Output:** $0.018/min\n\nSince we don't want to build our own media streaming infrastructure from scratch, it's worth noting that these models are already integrated with partners like LiveKit, Vercel,\n\n[LangChain](https://promptcube3.com/en/tags/langchain/), Agora, Pipecat, Fishjam, and Vision Agents.\n\nIf you're testing this out, you can find the models in Google AI Studio or via the Live API. The jump from the previous live models to 3.8 is a pretty significant step-change in how these agents feel and react.\n\n[Next OpenAI's software factory allows some pull requests to skip human review →](https://promptcube3.com/en/threads/9478/)\n\n## All Replies （3）\n\nExcited to see this finally working! I'm curious if the 15-20% jump happens with Whisper v3 or a different model?\n\nCurious if that asterisk implies a hidden tier or volume discount. I'm wondering if text output is just billed as standard tokens?\n\nI want to try this tonight. Does it actually work for MetaMask wallets or just CEX accounts?", "url": "https://wpnews.pro/news/gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent", "canonical_source": "https://promptcube3.com/en/threads/9494/", "published_at": "2026-09-17 16:45:19+00:00", "updated_at": "2026-09-17 16:52:49.246389+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-agents", "natural-language-processing", "ai-tools"], "entities": ["Google", "Gemini 3.8 Live", "Gemini 3.8 Live Extended Thinking", "Google AI Studio", "Gemini 3.5 Transcribe", "Artificial Analysis", "LiveKit", "LangChain"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent", "markdown": "https://wpnews.pro/news/gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent.md", "text": "https://wpnews.pro/news/gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent.txt", "jsonld": "https://wpnews.pro/news/gemini-3-8-live-is-finally-in-the-api-and-the-pricing-is-actually-decent.jsonld"}}