{"slug": "build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe", "title": "Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe", "summary": "Google released new Gemini Live models in the Gemini API and Google AI Studio, including Gemini 3.8 Live and 3.8 Live Extended Thinking, which the company says rank first on Artificial Analysis' Speech-to-Speech leaderboard. The company also highlighted Gemini 3.5 Transcribe, a speech-to-text model covering 85+ languages with an average Word Error Rate of 4.0% streaming and 2.6% non-streaming. The Live models are priced at $0.005/min for audio input and $0.018/min for audio output, with integrations through partners including Agora, LangChain, LiveKit, Pipecat, and Vercel.", "body_md": "Yesterday, we [released](https://blog.google/innovation-and-ai/models-and-research/gemini-models/new-gemini-dialogue-models/) new Gemini Live models in the [Gemini API](https://ai.google.dev/gemini-api/docs/live) and [Google AI Studio](http://ai.studio/live), expanding our developer suite for building real-time, voice-first product experiences: \n\n[**Gemini 3.8 Live**](https://aistudio.google.com/live?model=gemini-3.8-live) and **[3.8 Live Extended Thinking](https://aistudio.google.com/live?model=gemini-3.8-live-extended-thinking):** Gemini 3.8 Live brings a step change to our native speech-to-speech models, capable of performing tasks while maintaining dialogue. For complex requests, 3.8 Live Extended Thinking delivers deeper reasoning, ranking #1 on Artificial Analysis’ Speech-to-Speech leaderboard.\n\n[**Gemini 3.5 Transcribe**](https://aistudio.google.com/live?model=gemini-3.5-transcribe-live): Our dedicated speech-to-text model brings highly precise transcription across 85+ languages. Released last month, it achieved an average Word Error Rate (WER) of 4.0% (streaming) and 2.6% (non-streaming).\n\nOur new models, [Gemini 3.8 Live](https://aistudio.google.com/live?model=gemini-3.8-live) and [3.8 Live Extended Thinking](https://aistudio.google.com/live?model=gemini-3.8-live-extended-thinking) enable developers to build voice agents that can reason and execute tasks while maintaining the flow of conversations. Key capabilities include: \n\n3.8 Live Extended Thinking also supports [configurable thinking](https://ai.google.dev/gemini-api/docs/live-api/thinking) to help handle complex, multi-step reasoning in the background, while responding or narrating its progress in the main conversation. These models represent a step-change from our previous live models and provide a more streamlined alternative to cascaded architectures.\n\nGemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available via the [Live API](https://ai.google.dev/gemini-api/docs/live). Competitively [priced](http://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-live) at $0.005/min for audio input and $0.018/min* for audio output, they allow developers to scale voice applications with industry-leading performance.\n\nDevelopers can also access the models through [Agora](https://docs.agora.io/en/ai/models/mllm/gemini), [Fishjam](https://docs.fishjam.io/tutorials/gemini-live-integration), [LangChain](https://docs.langchain.com/langsmith/trace-gemini-live), [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gemini/), [Pipecat](https://docs.pipecat.ai/pipecat/features/gemini-live), [Vercel](https://vercel.com/docs/ai-gateway/modalities/realtime), and [Vision Agents](https://visionagents.ai/integrations/realtime/gemini), our Live API integration partners that handle media streaming infrastructure for real-world deployment.\n\nReal-time speech understanding is critical for voice-first interfaces. Last month, we released Gemini 3.5 Transcribe for low-latency transcription with high precision, achieving a 4.0% WER, and useful features:\n\n`custom_vocabulary` list of up to 1,000 terms\n3.5 Transcribe supports 85+ languages and provides a strong listening engine for voice experiences and stateless tasks like sub-second captioning, call center agents, and real-time audio analytics. You can also access the model via the Interactions API to transcribe audio files up to 1 hour long with structured timestamps and speaker labeling. Read our [developer guide](https://aistudio.google.com/learn/gemini-3-5-transcribe-developer-guide) to learn more.\n\nTo get started, try out the models in [ai.studio/live](https://ai.studio/live), clone example apps from [GitHub](https://github.com/google-gemini/gemini-live-api-examples), or equip your agent with our [live api skill](https://ai.google.dev/gemini-api/docs/coding-agents#gemini-live-api-dev).\n\nYou can also create audio experiences with our speech and music generation models, all available in the Gemini API:\n\nThe mic is yours, and we can’t wait to hear what you build!", "url": "https://wpnews.pro/news/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe", "canonical_source": "https://dev.to/googleai/build-real-time-voice-applications-with-gemini-38-live-and-35-transcribe-4nb5", "published_at": "2026-09-16 14:30:05+00:00", "updated_at": "2026-09-16 14:43:30.258518+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "natural-language-processing", "developer-tools"], "entities": ["Google", "Gemini 3.8 Live", "Gemini 3.8 Live Extended Thinking", "Gemini 3.5 Transcribe", "Gemini API", "Google AI Studio", "Artificial Analysis", "LiveKit"], "alternates": {"html": "https://wpnews.pro/news/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe", "markdown": "https://wpnews.pro/news/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe.md", "text": "https://wpnews.pro/news/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe.txt", "jsonld": "https://wpnews.pro/news/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5-transcribe.jsonld"}}