{"slug": "gemini-3-8-live-or-extended-thinking-in-famulor", "title": "Gemini 3.8 Live or Extended Thinking in Famulor?", "summary": "Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, two native speech-to-speech models that process speech directly and execute tools asynchronously while a conversation continues. According to a Famulor changelog dated September 17, 2026, both models are now part of Famulor's Realtime model selection, with manual model selection visible only when the Fallbacks & Guardrails add-on is included in the active workspace; otherwise Famulor uses its recommended automatic selection. Google positions Gemini 3.8 Live as the default for low-latency voice-agent experiences and Extended Thinking for complex, multi-step tasks requiring stronger background reasoning.", "body_md": "### Summarize Content With:\n\nOn September 15, 2026, Google introduced two new speech-to-speech models: **Gemini 3.8 Live** for fluid, low-latency dialogue and **Gemini 3.8 Live Extended Thinking** for more complex, multi-step tasks. According to the [Famulor changelog dated September 17, 2026](https://docs.famulor.io/changelog#2026-09-17), both models are now part of Famulor's Realtime model selection.\n\nThe useful question is not which model is universally “better.” It is: **Which kind of delay and failure is less acceptable in your call?** Extra reasoning may unnecessarily slow a short [appointment](https://www.famulor.io/use-cases/appointment-booking-faqs) qualification flow. In a multi-step rescheduling process with rules and tool calls, too little reasoning may make process errors more likely.\n\n**Key takeaways**\n\n- Gemini 3.8 Live is the starting point for fast, frequent conversation steps.\n- Extended Thinking belongs in tests with genuinely multi-step tasks, not in every call by default.\n- Famulor says manual model selection is visible with the\n**Fallbacks & Guardrails** add-on; otherwise Famulor uses its recommended automatic selection.- Decide with identical tests over the real phone path, not provider benchmarks alone.\n\n## What is new in Gemini 3.8 Live?\n\nGoogle's [September 15, 2026 announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/) describes both as native Live models that process speech directly and can execute tools asynchronously while the conversation continues. Google positions the standard model for scale and fluid dialogue, while Extended Thinking adds more background reasoning for complex tasks.\n\nThe technical model pages draw a clearer boundary:\n\n- [Gemini 3.8 Live](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live) is Google's default recommendation for low-latency voice-agent experiences without reasoning-induced delays. It supports asynchronous function calling and interleaved reasoning.\n- [Gemini 3.8 Live Extended Thinking](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live-extended-thinking) targets complex, multi-step tasks that need stronger background reasoning. Google also notes that asynchronous processing can continue after what appears to be the end of a conversational turn.\n\nThese are **Google product descriptions**, not independent performance evidence for your phone use case. Google's published benchmark results were not measured with your phone number, prompt, tools or caller profiles. They can shape a test hypothesis, but they should not replace a production acceptance test.\n\n## How are the models exposed in Famulor?\n\nFamulor describes Realtime as an engine mode in which one model listens and speaks. According to the current [Models & Voices documentation](https://docs.famulor.io/assistants/models-and-voices), Famulor selects the models used by default. A compatible model catalog becomes visible, and manual selection becomes available, when **Fallbacks & Guardrails** is included in the active workspace.\n\nThe Gemini 3.8 changelog entry highlights four practical points:\n\n1. Both 3.8 Live variants are part of the Realtime selection.\n2. Without the add-on, Famulor retains its recommended automatic selection.\n3. Extended Thinking can help with more demanding questions but may feel slower.\n4. Realtime improvements also cover tool execution, complete word endings and silent stretches of audio.\n\nAlways verify actual availability in the active workspace under **Settings → Plan**. This article does not promise a specific plan entitlement or permanent model availability.\n\n## When is Gemini 3.8 Live the better starting point?\n\nStart with Gemini 3.8 Live when the call consists of short, repetitive decisions and conversational rhythm matters most. Typical examples include:\n\n- opening hours, status checks and simple FAQs;\n- lead qualification with a few clear criteria;\n- appointment requests with a small set of required fields;\n- routing by intent, language or location;\n- short tool calls whose results can be confirmed immediately.\n\n“Fast” should not be turned into an invented millisecond threshold. What matters is whether the experience feels natural to callers: Does the assistant acknowledge the request early enough, avoid talking over the caller and remain clear after a tool call?\n\nThe standard model is also sensible when complex cases are handed to a human or a specialized assistant. In that design, every simple call does not have to carry the cost of deeper reasoning.\n\n## When should you test Extended Thinking?\n\nExtended Thinking deserves a separate test path when several steps depend on each other and a premature intermediate decision would be costly or difficult to correct. Examples include:\n\n- rescheduling that combines availability, rate rules and customer constraints;\n- a service case with diagnostic questions, knowledge retrieval and ticket creation;\n- a win-back offer with exclusions and approval rules;\n- a booking flow with several asynchronous tool calls;\n- a process in which the assistant should explain progress while background work continues.\n\nMore reasoning is not permission for unlimited autonomy. Continue to define allowed tools, mandatory confirmations, stop conditions and handoffs explicitly. Payments, contract changes, medical information and other consequential actions should retain appropriate confirmation or human approval in the workflow.\n\n## Decision matrix: conversational flow or reasoning depth?\n\n| Question | Lean toward Gemini 3.8 Live | Test Extended Thinking | \n|---|---|---|\n| How many dependent steps make up the core process? | few | several | \n| Must the assistant keep speaking while a tool runs? | a short acknowledgement is enough | progress and intermediate steps matter | \n| How harmful is an extra conversational pause? | highly disruptive | acceptable if it improves process reliability | \n| How hard is a wrong intermediate decision to correct? | easy | difficult | \n| Is there a clear escalation route? | yes, early | yes, after structured checks | \n| What is the primary success criterion? | fluid completion of a simple task | correct completion of a complex task | \n\nThis matrix is a starting point, not a vendor ranking. If your scenario lands in both columns, split the flow: route simple requests through the fast path and complex requests through a bounded path with more reasoning or a human handoff.\n\n## How to run a fair A/B test in Famulor\n\n### 1. Keep the conversation design constant\n\nUse the same system prompt, tools, knowledge sources, voice and test intents for both variants. Change only the model selection. Otherwise, you cannot reliably attribute an improvement.\n\nThe current [Famulor engine-mode overview](https://docs.famulor.io/assistants/engine-modes) also marks an important boundary: Realtime exposes fewer component-level controls than a pipeline composed of STT, LLM and TTS. If you need a specific brand or cloned voice, first confirm that Realtime is the right architecture. Famulor already has a separate [Realtime versus pipeline architecture guide](https://www.famulor.io/en/blog/realtime-vs-pipeline-voice-agent-architecture-guide-2026); this article focuses only on choosing between the new Gemini 3.8 Live models.\n\n### 2. Define a small, demanding test suite\n\nDo not test only the happy path. A useful suite includes at least:\n\n- a simple request without a tool;\n- a successful tool call;\n- a slow or failed tool call;\n- a caller correction in the middle of the process;\n- an ambiguous detail;\n- a request outside allowed actions;\n- a human handoff;\n- a repeat call with a similar but non-identical intent.\n\nScore task completion, captured-data accuracy, tool sequence, conversational flow, interruption handling and handoff quality separately. Avoid collapsing them into a single average too early.\n\n### 3. Test over the real phone path\n\nBrowser demos and provider recordings do not show how your assistant behaves across phone numbers, networks and real devices. Repeat the same cases with typical mobile phones, background noise, speaking rates and accents. Log observable failures such as “tool fired twice,” “correction ignored” or “long silence before acknowledgement.”\n\n### 4. Set stop and acceptance rules in advance\n\nA model should not win because one impressive call sounds unusually good. Define P0 failures before testing, such as unconfirmed consequential actions, wrong tool parameters, missing required data or a failed handoff. A single P0 should not be offset by fluid prosody.\n\n## Practical example: service appointment rescheduling\n\nA service company wants to automate inbound rescheduling calls. The assistant must capture the work-order number, retrieve the existing appointment, check available slots, consider technician coverage and create the new booking only after explicit confirmation.\n\nFor the straightforward case—a valid order number, one available slot and immediate confirmation—Gemini 3.8 Live is a plausible starting point. The conversation should stay short and the tool result should be confirmed directly.\n\nFor the difficult case—multiple orders, an ambiguous address, two calendar lookups and a rate rule—the team tests Extended Thinking. The assistant can announce intermediate steps, but it must summarize the date, time window and order before making a change. If a rule cannot be resolved confidently, it hands the call over instead of guessing.\n\nThis test does not assume that Extended Thinking automatically completes more reschedules correctly. It checks whether the added reasoning architecture produces measurably fewer process failures **in this defined workflow** without slowing the conversation beyond an acceptable level.\n\n## What does the practitioner community signal?\n\n**Community signal, not product evidence:** In a [Show HN post dated September 10, 2026](https://news.ycombinator.com/item?id=49646928), voice-AI developers say repeated manual test calls and unexpected production cases motivated them to build a simulation platform. The thread had limited engagement when reviewed. It proves neither the project's effectiveness nor any model's quality, but it supports a useful testing question: Is one good demo call enough, or is the workflow stable across many reproducible scenarios?\n\nThe distinction is deliberate: Google and Famulor support claims about product features and availability. Hacker News provides a practitioner signal about an operational concern. It cannot replace official documentation or your own evaluation.\n\n## Which limits should you document before rollout?\n\n- **No benchmark transfer:** Google's scores are not automatically your production results.\n- **No availability guarantee:** Plans, catalogs and automatic selection can change; check the active workspace.\n- **No automatic compliance:** Model choice does not establish legal permissibility or replace privacy, recording and consent reviews.\n- **No unlimited autonomy:** Tool permissions, confirmations and handoffs remain design decisions.\n- **No one-time test:** Changes to prompts, tools, knowledge or models require regression testing.\n- **No substitute for architecture selection:** If a specific TTS voice or component-level control is mandatory, pipeline or half-cascade may be more suitable.\n\n## Conclusion: Choose by task profile, not model name\n\nGemini 3.8 Live is the logical starting point for fluid, frequent realtime conversations. Extended Thinking is a targeted option when multi-step reasoning and asynchronous tools may deliver a real process benefit. “More thinking” is not a universal quality upgrade; it must be weighed against conversational flow, failure modes and escalation needs.\n\nThe cleanest decision comes from a controlled comparison with the same prompt, tools and reproducible phone scenarios. Only a variant that passes your predefined quality gates in your own workflow should reach production traffic.\n\n### Test both Realtime variants with the same call suite\n\nFirst check whether manual model selection is available in your Famulor workspace. Then build one simple and one complex path and compare both variants through real test calls. [Start with Famulor's Realtime documentation](https://docs.famulor.io/assistants/engine-modes#realtime).\n\n## FAQ about Gemini 3.8 Live in Famulor\n\n### Can every Famulor workspace select Gemini 3.8 Live manually?\n\nNo. Famulor says model selection is visible when **Fallbacks & Guardrails** is included in the active workspace. Without the add-on, Famulor uses its recommended automatic selection. Check the current plan in your workspace.\n\n### Is Extended Thinking always more accurate?\n\nThat cannot be claimed universally. Google positions it for complex, multi-step tasks. Whether it makes your workflow more reliable must be tested with identical scenarios, tools and acceptance criteria.\n\n### Which model should I use for simple appointment qualification?\n\nGemini 3.8 Live is the natural starting point when you capture a few fields and confirm short tool calls. Still test interruptions, corrections and tool failures over the real phone path.\n\n### When should I test Extended Thinking?\n\nWhen several rules, lookups or tools depend on each other and a wrong intermediate decision would be hard to correct. Continue to define clear confirmations, boundaries and handoffs.\n\n### Does Realtime replace an STT-LLM-TTS pipeline?\n\nNot for every use case. Realtime provides an integrated conversation path with fewer component-level controls. If you must control components or a specific TTS voice separately, evaluate pipeline or half-cascade.\n\n### Is one successful test call enough for rollout?\n\nNo. Use reproducible normal, failure and edge cases, test real phone conditions and repeat the suite after relevant changes.\n\nWriter at Famulor", "url": "https://wpnews.pro/news/gemini-3-8-live-or-extended-thinking-in-famulor", "canonical_source": "https://www.famulor.io/blog/gemini-38-live-vs-extended-thinking-famulor", "published_at": "2026-09-20 15:23:00+00:00", "updated_at": "2026-09-21 16:25:21.435375+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-agents", "natural-language-processing"], "entities": ["Google", "Gemini 3.8 Live", "Gemini 3.8 Live Extended Thinking", "Famulor", "Fallbacks & Guardrails"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-8-live-or-extended-thinking-in-famulor", "markdown": "https://wpnews.pro/news/gemini-3-8-live-or-extended-thinking-in-famulor.md", "text": "https://wpnews.pro/news/gemini-3-8-live-or-extended-thinking-in-famulor.txt", "jsonld": "https://wpnews.pro/news/gemini-3-8-live-or-extended-thinking-in-famulor.jsonld"}}