{"slug": "vibevoice-asr-streaming-technical-report", "title": "VibeVoice-ASR-Streaming Technical Report", "summary": "Researchers released VibeVoice-ASR-Streaming, an LLM-based end-to-end streaming speaker-attributed automatic speech recognition system that processes fixed-size audio chunks with lookahead to output 'who said what' in real time without a separate diarization stage. The 7B model achieves the lowest average WER/CER across five evaluation sets and best or tied-best speaker attribution on 12 of 13 settings; the team released 1.5B and 7B model weights with inference code.", "body_md": "# Electrical Engineering and Systems Science > Audio and Speech Processing\n\n[Submitted on 2 Sep 2026]\n\n# Title:VibeVoice-ASR-Streaming Technical Report\n\n[View PDF](/pdf/2609.02812v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.02812v1)\n\nAbstract:Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/vibevoice-asr-streaming-technical-report", "canonical_source": "http://arxiv.org/abs/2609.02812v1", "published_at": "2026-09-03 13:47:48+00:00", "updated_at": "2026-09-03 14:22:42.437613+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "natural-language-processing", "ai-research"], "entities": ["VibeVoice-ASR-Streaming", "VibeVoice-ASR", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/vibevoice-asr-streaming-technical-report", "markdown": "https://wpnews.pro/news/vibevoice-asr-streaming-technical-report.md", "text": "https://wpnews.pro/news/vibevoice-asr-streaming-technical-report.txt", "jsonld": "https://wpnews.pro/news/vibevoice-asr-streaming-technical-report.jsonld"}}