Tool: Gemini Live audio Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.
I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking.
The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.
Here's [the Gemini Live tutorial](https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket) for getting started with that WebSockets API.
Tags: [google](https://simonwillison.net/tags/google), [tools](https://simonwillison.net/tags/tools), [websockets](https://simonwillison.net/tags/websockets), [generative-ai](https://simonwillison.net/tags/generative-ai), [llms](https://simonwillison.net/tags/llms), [gemini](https://simonwillison.net/tags/gemini), [llm-release](https://simonwillison.net/tags/llm-release), [speech-to-text](https://simonwillison.net/tags/speech-to-text)