{"slug": "ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test", "title": "[AI in Action] Gemini 3.8 Live Upgrade: Paying Off Technical Debt and How My Test Script Fooled Me Three Times", "summary": "A developer upgraded their LINE Bot's LIFF voice assistant from the preview-only Gemini 3.1 Flash Live model to Google's newly released Gemini 3.8 Live, removing a long-standing preview-model exemption in their test guards. After querying the API directly to confirm available model IDs, they found gemini-3.8-live has identical token limits and generation methods to its predecessor, though they caution the swap is not yet proven behaviorally painless.", "body_md": "I have a LINE Bot that I use every day, [linebot-helper-python](https://github.com/kkdai/linebot-helper-python). It provides summaries and social media copy for URLs, allows continuous questioning for YouTube links, and handles bookmarks and location queries. One of its features is a LIFF voice assistant: it opens a webpage, connects to the Gemini Live API via WebSocket, and supports push-to-talk or hands-free conversation. The transcribed content is then summarized into a message and pushed back to the LINE chatroom.\n\nIn mid-September, Google released [Gemini 3.8 Live](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/). After reading the announcement, what I wanted to do was slightly different from what they were trying to sell.\n\nThe announcement focused on capabilities: Speech to Speech Index ranked first (82.6 points), Big Bench Audio at 97.7%, automatic detection of 97 languages with the ability to switch mid-conversation, tool calling executing in the background without interrupting the flow, and real-time visual input. There is also a `Gemini 3.8 Live Extended Thinking` model that can reason while speaking.\n\nMy first thought was: **Now I can finally remove that preview exception.**\n\nIn my previous post about agentic video, I mentioned that my intelligent dialogue feature was broken for a while, and no one noticed. The reason was that `loader/chat_session.py` had `gemini-3-pro-preview` hardcoded. After that model was removed from Vertex AI, the actual path returned a 404, and the exception handling in that file simply raised the error.\n\nAfter fixing it, I added a set of guards using tests to enforce that \"Model IDs can only appear in `config/agent_config.py`\" and \"No preview or experimental models allowed.\" But those guards had two exceptions:\n\n```\n# Live API only supports gemini-3.1-flash-live-preview; TTS is a separate model family.\n# These two places must maintain preview models and are outside the scope of the guards.\nEXEMPT_FILES = {\"services/voice_live.py\", \"tools/tts_tool.py\"}\n```\n\nThe reason for the voice exception was \"The Live API only has this one model available.\" At the time, that was true. But once that sentence is written into a comment, no one checks if it's still valid—the last loophole for the same type of bug sat there in plain sight for months.\n\nThe significance of 3.8 Live, therefore, isn't just the capability upgrade. Its model name **does not contain `-preview`**, which means that exception can be removed.\n\nThis time, I didn't read the documentation first. Instead, I asked the API directly what actually exists:\n\n```\ncurl -s \"https://generativelanguage.googleapis.com/v1beta/models?key=$KEY&pageSize=200\" \\\n  | jq -r '.models[]?.name' | rg -i 'live'\n\nmodels/gemini-3.5-transcribe-live\nmodels/gemini-3.1-flash-live-preview\nmodels/gemini-3.8-live\nmodels/gemini-3.8-live-extended-thinking\nmodels/gemini-3.5-live-translate-preview\n```\n\nOne command gave me five confirmed model names, which is much more reliable than digging through documentation. `gemini-3.8-live` is right there, with a clean name.\n\nI also happened to find `gemini-3.5-live-translate-preview`, which I didn't know about. I was originally thinking about \"whether to implement a real-time interpretation mode,\" and it seems that's a specialized model rather than a brute-forced general model. I'll look into that later.\n\nNext, I queried the metadata:\n\n|  | `3.1-flash-live-preview` | `3.8-live` | \n|---|---|---|\n| inputTokenLimit | 131072 | 131072 | \n| outputTokenLimit | 65536 | 65536 | \n| supportedGenerationMethods | `bidiGenerateContent` | `bidiGenerateContent` | \n\nThe numbers are identical. At this point, I had initial confidence that it would be a \"painless swap,\" but this was just metadata, not behavior.\n\nWhen I first evaluated this, I solemnly noted: \"The announcement didn't mention Vertex AI, and the decision for this project is to switch entirely to Vertex, so it's uncertain if 3.8 Live will be available.\"\n\nIt sounded very professional. Then I looked at my own code:\n\n```\nclient = live_genai.Client(\n    api_key=GOOGLE_AI_API_KEY,\n    vertexai=False,\n    http_options={\"api_version\": \"v1beta\"},\n)\n```\n\n`vertexai=False`. The voice path has used the Gemini API with an AI Studio key from the very beginning; it doesn't use Vertex at all—and 3.8 Live was released on the Gemini API.\n\nI was warning my own repo about a risk that didn't exist in my repo. The project as a whole indeed \"switched entirely to Vertex AI,\" but voice was the exception, and I was the one who wrote that exception.\n\n**Cause and Solution**: Project-level decision records can become a memory shortcut. \"We use Vertex for everything\" is a correct summary, but summaries erase exceptions. Checking the source code takes thirty seconds and is much cheaper than reasoning based on impressions.\n\nAfter confirming the model existed, I wrote a throwaway script for testing. The key design was: **Don't write a separate config; directly import the existing `build_live_config()` and `build_voice_tools()` from the project.** I wanted to test \"Can the configuration I'm currently running work?\", not \"Can a separate configuration I wrote work?\".\n\nThe first run looked like this:\n\n```\n=== gemini-3.8-live ===\n  [handsfree] {'connected': True, 'audio': True, ..., 'error': None}\n  [PTT] {'connected': True, 'audio': False, ...,\n               'error': 'APIError: 1007 None. Precondition check failed.'}\n```\n\nThe push-to-talk (PTT) mode was blocked. If I had only tested 3.8, I would have concluded: \"3.8 Live doesn't support PTT; cannot upgrade.\"\n\nBut I also ran the old model:\n\n```\n=== gemini-3.1-flash-live-preview ===\n  [PTT] {..., 'error': 'APIError: 1007 None. Precondition check failed.'}\n```\n\nExactly the same error. Since both models failed in the same way, it wasn't a model issue; it was a script issue.\n\nThe actual error was: In PTT mode, the production environment sends PCM audio, but I sent text for convenience. Inserting text input between `activity_start` and `activity_end` is an invalid combination.\n\nAfter switching to real 16kHz PCM:\n\n```\n=== gemini-3.8-live ===\n  [PTT] {'connected': True, 'audio': True, 'out_tx': True, 'in_tx': True,\n         'resume': True, 'events': ['audio','in_tx','out_tx','resume','turn_complete']}\n```\n\nEverything was there, and it matched the old model item by item.\n\n**Cause and Solution**: The only thing I did right this time was **keeping a control group**. When testing something new, testing the old thing in the same way costs almost nothing, but it allows you to distinguish between \"the new thing is broken\" and \"my test is written incorrectly.\" These two things look identical.\n\nIn the second version of the script, I used `math.sin` to generate a 180Hz PCM segment as audio input. PTT mode tested smoothly, but both models timed out in hands-free mode.\n\nI almost wrote this off as \"hands-free behavior pending confirmation.\" But a moment's thought revealed why: hands-free mode relies on Gemini's own Voice Activity Detection (VAD), and a pure sine wave isn't speech. The VAD correctly judged that \"no one is talking,\" so it never triggered.\n\nSo it wasn't a \"detected problem\"; it was **not detected at all**. Two timeouts looked like data, but they were actually two empty spaces.\n\nThis later evolved into a checklist. I separated what the script tested from what it didn't:\n\nTested: Connection, config acceptance, PTT audio round-trip, bidirectional transcripts, resumption handle, turn_complete, tool calling trigger with correct parameters.\n\n**Not tested**: Real human Chinese speech recognition, automatic VAD in hands-free mode, interrupting while speaking, reconnection after a ten-minute connection recycle, whether the resumption handle actually restores context, Google Search grounding actually taking effect, the second half of slow task delegation pushed back to LINE, voice quality, and latency.\n\nOf the eight items, eight were not tested. The script ran beautifully, but it touched the protocol layer, not what a user would encounter.\n\nAfter going through the process, the final comparison looks like this. Both columns are results from hitting the real API using the **project's existing config generator**:\n\n| Validation Item | `3.1-flash-live-preview` | `3.8-live` | \n|---|---|---|\n| PTT (activity signal + PCM chunks) | Audio / Bidirectional Transcripts / resume / turn_complete | Same | \n| `session_resumption` | Yes | Yes | \n| `context_window_compression` | Accepted | Accepted | \n| Bidirectional `AudioTranscriptionConfig` | Yes | Yes | \n| Voice `Aoede` | Yes | Yes | \n| `google_search` and`function_declarations` in separate Tools | Accepted | Accepted | \n| Tool calling actually triggered | Both tool parameters correct | Same | \n| `api_version=\"v1beta\"` | Required | Still required | \n\nIn other words, `gemini-3.8-live` can be swapped directly into the existing configuration.\n\nI also tested `gemini-3.8-live-extended-thinking`, and the connection was immediately blocked:\n\n```\nAPIError: 1007 None. Thinking level must be specified for this model.\n```\n\nAfter adding `thinking_config`, both `LOW` and `HIGH` could connect. So it's not broken; it just has an extra required parameter.\n\nI decided not to adopt it because the current `build_live_config()` doesn't send `thinking_config`. If I just swapped the model name, the voice assistant would crash at the connection step. Using it would require modifying the config generator, which is a separate task.\n\nBut I couldn't \"keep\" the knowledge that it shouldn't be set, so I wrote it as a test:\n\n```\n# Live API (bidiGenerateContent) support list. Confirmed by connection test on 2026-09-16.\n# Deliberately excludes gemini-3.8-live-extended-thinking: this model mandates thinking_level,\n# and connections without thinking_config will be blocked (1007).\nLIVE_CAPABLE = {\"gemini-3.8-live\"}\n```\n\nLast time I wrote \"Live API only supports a certain model\" in a comment, that sentence expired for months without anyone noticing. This time, I wrote it as an assertion that will fail.\n\nThe actual code changes were small, about eighty lines across seven files:\n\n`config/agent_config.py`` VOICE_MODEL` module constant and `AgentConfig.voice_model`. A detail here—I deliberately didn't just put it in `get_agent_config()`, because that function requires `GOOGLE_CLOUD_PROJECT`, and `services/voice_live.py` needs the model ID at import time. If I went that route, imports would fail if the project environment variable wasn't set.`tests/test_model_config.py`` EXEMPT_FILES` from two files to one; included `voice_model` in the existing \"no preview\" and \"must be in the verified capable list\" guards.\nThe last guard originally looked like this:\n\n``` python\ndef test_voice_live_model_unchanged():\n    \"\"\"Live API only supports gemini-3.1-flash-live-preview; changing it will break the voice assistant.\"\"\"\n    assert VOICE_MODEL == \"gemini-3.1-flash-live-preview\"\n```\n\nIt asserted a string. Strings expire, and when they do, the test is still green—it only gets in your way when you want to upgrade; it doesn't save you when a model is taken down.\n\nBy changing it to a list of capable models, it guards against the category of \"this value must be a verified Live model,\" rather than a specific value.\n\nThe exception for `tools/tts_tool.py` remains. I checked, and all three TTS models on the API (`2.5-flash-preview-tts`, `2.5-pro-preview-tts`, `3.1-flash-tts-preview`) are still in preview, with no GA versions to swap to. The exceptions were reduced from two to one, but not zeroed out.\n\nEnvironment variables were set before merging:\n\n```\ngcloud run services update linebot-helper-python --region us-central1 \\\n  --update-env-vars VOICE_MODEL=gemini-3.8-live\n```\n\nThis step had **absolutely no effect** at the time. The live version was still running the old image, which had the model hardcoded in `voice_live.py` and wouldn't even read this variable.\n\nBut the order was correct. Merging triggers Cloud Build for automatic deployment. As soon as the new image goes up, it reads the already-in-place configuration, ensuring there isn't a window where \"the code is up but the config isn't.\"\n\nThis approach has a side effect that I find more valuable than the upgrade itself: **Rollback doesn't require re-deployment.**\n\n```\ngcloud run services update linebot-helper-python --region us-central1 \\\n  --update-env-vars VOICE_MODEL=gemini-3.1-flash-live-preview\n```\n\nIf real-world testing feels off, a single command switches it back without waiting for a build. Changing the model ID from hardcoded to an environment variable bought me this flexibility.\n\nThere was one thing in the script I always found suspicious. During the tool calling round, the old model would speak a filler phrase before waiting for the tool result, while 3.8 did not.\n\nMy judgment at the time was cautious: this **might** be the \"tools execute in the background without interrupting the flow\" mentioned in the announcement, or it might just be my loop closing early. Automation couldn't tell; it required a human talking to know if it was an improvement or an awkward silence.\n\nAfter merging and deploying, I spoke a few sentences via LIFF. The answer: **There is no difference from before; it's very smooth.**\n\nSo that difference doesn't exist to a human ear. It was a byproduct of my loop returning as soon as it got `turn_complete`. I was observing the shape of my own script.\n\nIn this article, my script lied to me three times: the PTT incident made the new model look broken, the sine wave incident made \"not tested\" look like \"tested,\" and this time made a non-existent difference look like a new capability. None of them were model issues.\n\nThis is the part I find most worth writing down.\n\nAfter the upgrade, the tests went from 293 to 296, all green. But this number has almost nothing to do with \"whether voice works on 3.8.\"\n\n`tests/test_voice_live.py` has 32 tests covering PTT disabling auto-VAD, activity signals, transcript forwarding, tool calling execution, resumption handles, connection recycling, and interruptions. It looks comprehensive. But they all use `FakeLiveSession` and don't connect to the internet. If the model changes from 3.1 to 3.8, not a single one of these 32 tests will change color.\n\nThey test \"Is what I'm sending correct?\", not \"What is the other side returning?\".\n\nThe three layers of coverage look like this:\n\n| Layer | What it tests | Who runs it | Repeatable | \n|---|---|---|---|\n| Repo test suite (296) | Outgoing config and message format | CI on every push | Yes | \n| Throwaway probe (3) | Protocol layer compatibility with real API | Ran only once | No | \n| Real human conversation | Understanding, smoothness | Me | No | \n\nThe middle layer is the only automated test that actually touched 3.8, and it's not in the repo, CI doesn't run it, and no one will know the next time a model is taken down.\n\nThe newly added guard only blocks \"setting it to a known bad value\"; it doesn't block \"Google taking 3.8-live down\"—which is what actually happened last time.\n\nThat loophole is still open. Turning the probe into a smoke test that requires a key is the solution; I haven't done it yet.\n\nThe code is at [kkdai/linebot-helper-python](https://github.com/kkdai/linebot-helper-python), and this change is in PR #24. The official announcement is [Gemini 3.8 Live](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/), which mentions that all AI-generated audio carries a SynthID watermark—I did not verify this.\n\nThe announcement spent the most space on real-time visual input, which I didn't touch at all this time—currently, the LIFF `getUserMedia` is hardcoded to `video: false`. That will probably be the next post.", "url": "https://wpnews.pro/news/ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test", "canonical_source": "https://dev.to/evanlin/ai-in-action-gemini-38-live-upgrade-paying-off-technical-debt-and-how-my-test-script-fooled-me-5dh3", "published_at": "2026-09-29 08:39:05+00:00", "updated_at": "2026-09-29 08:46:37.738596+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "large-language-models", "developer-tools"], "entities": ["Google", "Gemini 3.8 Live", "Gemini 3.1 Flash Live", "Vertex AI", "LINE", "linebot-helper-python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test", "markdown": "https://wpnews.pro/news/ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test.md", "text": "https://wpnews.pro/news/ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test.txt", "jsonld": "https://wpnews.pro/news/ai-in-action-gemini-3-8-live-upgrade-paying-off-technical-debt-and-how-my-test.jsonld"}}