# Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

> Source: <https://dev.to/googleai/build-real-time-voice-applications-with-gemini-38-live-and-35-transcribe-4nb5>
> Published: 2026-09-16 14:30:05+00:00

Yesterday, we [released](https://blog.google/innovation-and-ai/models-and-research/gemini-models/new-gemini-dialogue-models/) new Gemini Live models in the [Gemini API](https://ai.google.dev/gemini-api/docs/live) and [Google AI Studio](http://ai.studio/live), expanding our developer suite for building real-time, voice-first product experiences: 

[**Gemini 3.8 Live**](https://aistudio.google.com/live?model=gemini-3.8-live) and **[3.8 Live Extended Thinking](https://aistudio.google.com/live?model=gemini-3.8-live-extended-thinking):** Gemini 3.8 Live brings a step change to our native speech-to-speech models, capable of performing tasks while maintaining dialogue. For complex requests, 3.8 Live Extended Thinking delivers deeper reasoning, ranking #1 on Artificial Analysis’ Speech-to-Speech leaderboard.

[**Gemini 3.5 Transcribe**](https://aistudio.google.com/live?model=gemini-3.5-transcribe-live): Our dedicated speech-to-text model brings highly precise transcription across 85+ languages. Released last month, it achieved an average Word Error Rate (WER) of 4.0% (streaming) and 2.6% (non-streaming).

Our new models, [Gemini 3.8 Live](https://aistudio.google.com/live?model=gemini-3.8-live) and [3.8 Live Extended Thinking](https://aistudio.google.com/live?model=gemini-3.8-live-extended-thinking) enable developers to build voice agents that can reason and execute tasks while maintaining the flow of conversations. Key capabilities include: 

3.8 Live Extended Thinking also supports [configurable thinking](https://ai.google.dev/gemini-api/docs/live-api/thinking) to help handle complex, multi-step reasoning in the background, while responding or narrating its progress in the main conversation. These models represent a step-change from our previous live models and provide a more streamlined alternative to cascaded architectures.

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available via the [Live API](https://ai.google.dev/gemini-api/docs/live). Competitively [priced](http://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-live) at $0.005/min for audio input and $0.018/min* for audio output, they allow developers to scale voice applications with industry-leading performance.

Developers can also access the models through [Agora](https://docs.agora.io/en/ai/models/mllm/gemini), [Fishjam](https://docs.fishjam.io/tutorials/gemini-live-integration), [LangChain](https://docs.langchain.com/langsmith/trace-gemini-live), [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gemini/), [Pipecat](https://docs.pipecat.ai/pipecat/features/gemini-live), [Vercel](https://vercel.com/docs/ai-gateway/modalities/realtime), and [Vision Agents](https://visionagents.ai/integrations/realtime/gemini), our Live API integration partners that handle media streaming infrastructure for real-world deployment.

Real-time speech understanding is critical for voice-first interfaces. Last month, we released Gemini 3.5 Transcribe for low-latency transcription with high precision, achieving a 4.0% WER, and useful features:

`custom_vocabulary` list of up to 1,000 terms
3.5 Transcribe supports 85+ languages and provides a strong listening engine for voice experiences and stateless tasks like sub-second captioning, call center agents, and real-time audio analytics. You can also access the model via the Interactions API to transcribe audio files up to 1 hour long with structured timestamps and speaker labeling. Read our [developer guide](https://aistudio.google.com/learn/gemini-3-5-transcribe-developer-guide) to learn more.

To get started, try out the models in [ai.studio/live](https://ai.studio/live), clone example apps from [GitHub](https://github.com/google-gemini/gemini-live-api-examples), or equip your agent with our [live api skill](https://ai.google.dev/gemini-api/docs/coding-agents#gemini-live-api-dev).

You can also create audio experiences with our speech and music generation models, all available in the Gemini API:

The mic is yours, and we can’t wait to hear what you build!
