ByteDance's Seed organization has launched SeedRealtime, a native audio-visual full-duplex large language model built to watch, listen, and speak over continuous streams at once. The model combines audio, video, and text into a single architecture, allowing perception, understanding, decision-making, and expression to run in parallel. This avoids handoffs between separate speech recognition, vision-language, and text-to-speech modules, which can add latency and cause context loss.
The central shift is that the model determines conversational timing rather than relying on external voice-activity detection rules. SeedRealtime tracks scenes, speakers, s, and background chatter to judge what matters and when to answer. Visual context can resolve homophones, connect words such as "this" to a gesture or an earlier action, and retain information that has moved off-screen. It can also speak without a new prompt when a requested object appears, or it detects a mistake.
ByteDance showed these abilities in noisy, open settings. SeedRealtime matched names, faces, and voices during a group dinner, interpreted dishes and speech in a restaurant, and issued a reminder when a requested museum object entered view. It corrected an espresso-making mistake, spotted a requested section while pages of a paper were turning, ignored unrelated airport chatter, and followed a child's pointing during an English lesson. In the airport example, it went online to provide baggage-carousel information.
In end-to-end human evaluation, ByteDance says SeedRealtime cut audio-visual conversational pacing problems by half against cascaded models. Evaluators found fewer cutoffs, slow replies after s, and false triggers from nearby speech, while more conversations were completed smoothly. The announcement does not disclose sample sizes, a benchmark table or an exact figure for the completion gain.
Published by ByteDance's Seed organization, the release is framed as a step toward omni-modal systems that can observe, converse and act in changing real-world settings. Seed says the model is fully rolled out, but does not specify an API, product surface, supported regions, pricing or access requirements. Its roadmap covers lower latency, finer timing for interruptions and backchannels, stronger speaker tracking in multi-person scenes, proactive decisions and tool-connected tasks such as lookups and bookings.