# ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

> Source: <https://www.marktechpost.com/2026/08/09/bytedance-seed-introduces-seedrealtime-a-native-audio-visual-full-duplex-llm-that-watches-listens-and-speaks-in-one-model/>
> Published: 2026-08-10 05:48:17+00:00

ByteDance’s Seed team has introduced [SeedRealtime](https://seed.bytedance.com/en/SeedRealtime), a native audio-visual full-duplex LLM. The model fuses audio, video and text in a single unified architecture. It interacts in real time over continuous multimodal streams, rather than one turn at a time. Seed positions it as a step toward omni-modal interaction, and claims three breakthroughs: joint audio-visual understanding, proactive interaction, and natural conversational timing. The architectural target is the cascade: chained ASR, VLM and TTS modules that add latency and lose information between stages. SeedRealtime instead runs perception, understanding, decision-making and expression in parallel inside one end-to-end model. Turn-taking moves inside the model as well, replacing the external voice-activity detector most real-time stacks still depend on.

**Is it deployable?**

**It is partly deployable.**

SeedRealtime is [live inside the Doubao app](https://technode.com/2026/08/05/bytedance-launches-seedrealtime-full-duplex-audio-video-model/), ByteDance’s consumer assistant. For this specific model, ByteDance has published no technical report, no parameter count, no open weights, and no [Volcano Engine](https://www.volcengine.com/) or [BytePlus](https://www.byteplus.com/) endpoint. As a third-party team, you cannot integrate it as of now. What is deployable right now is the *idea*: a validated reference architecture, and a moved goalpost for anyone shipping real-time voice-plus-camera products.

**Interactive explainer**

**What is actually new in the demos**

Seed published seven scenarios. Four are load-bearing.

**Identity binding across modalities**: At a noisy group dinner, the model matches names to faces as people are introduced, then keeps each voice tied to its identity — attributing conflicting travel preferences to the right speaker before proposing a plan.**Proactive speech from a held instruction**: At the Hebei Museum, a user asks to be reminded when a specific bronze screen stand appears. The camera keeps panning; the model watches and speaks up unprompted when the piece enters frame. The same behavior shows up on a ResNet paper — the model tracks fast page flips, spots the “3.4 Implementation” section, pauses on its own, and reads out learning rate, momentum and weight decay.**Correction from visual state, not from a question**: Watching an espresso workflow, the model interrupts when whole beans go into the portafilter, then reads crema color and volume and suggests shortening extraction by 2 to 3 seconds.**Interference suppression pand off-screen memory**: At Beijing Daxing Airport, unrelated chatter about a flight does not trigger a reply. When the user actually asks, the model answers from departure-board information that has already scrolled off screen, and goes online for the baggage-carousel location.

**Key Takeaways**

- SeedRealtime is a native audio-visual full-duplex LLM — audio, video and text in one end-to-end architecture.
- Turn-taking moves inside the model; no external VAD decides when to speak.
- ByteDance’s own human eval reports pacing issues halved versus cascaded stacks — no benchmark, no latency numbers.
- It is live in the Doubao app, but there is no technical report, no weights and no announced API.

**Check out the **[ ByteDance Seed launch post](https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction) and

[.](https://seed.bytedance.com/en/models)

**Seed models page****Also, feel free to follow us on**

**and don’t forget to join our**[Twitter](https://x.com/intent/follow?screen_name=marktechpost)

**and Subscribe to**

[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)**. Wait! are you on telegram?**

[our Newsletter](https://www.aidevsignals.com/)

[now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? [Connect with us](https://forms.gle/wbash1wF6efRj8G58)

Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.
