{"slug": "dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis", "title": "DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis", "summary": "Researchers introduced DLLM-TTS, a text-to-speech framework using block discrete diffusion over X-Codec2 neural audio codec tokens, achieving a real-time factor of 0.15 with a 0.6B-parameter model trained on 20K hours, and competitive performance on the Seed-TTS-eval benchmark.", "body_md": "arXiv:2608.00011v1 Announce Type: new\nAbstract: Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and decode tokens sequentially, while non-autoregressive approaches im- prove speed at the cost of linguistic accuracy. We present DLLM-TTS, a framework that formulates TTS as conditional block discrete diffusion over X-Codec2 neural audio codec to- kens. The model decomposes sequences into blocks and applies masked diffusion within each block while processing blocks se- quentially, learning both local acoustic coherence and global text-speech alignment. During inference, parallel token pre- diction within blocks enables efficient generation with a real- time factor (RTF) of 0.15. A 0.6B-parameter model trained on 20K hours achieves competitive performance on the Seed- TTS-eval benchmark, demonstrating that block discrete diffu- sion language models enable practical and data-efficient speech synthesis with parallel generation.", "url": "https://wpnews.pro/news/dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis", "canonical_source": "https://arxiv.org/abs/2608.00011", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 04:35:34.490406+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "natural-language-processing"], "entities": ["DLLM-TTS", "X-Codec2", "Seed-TTS-eval"], "alternates": {"html": "https://wpnews.pro/news/dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis", "markdown": "https://wpnews.pro/news/dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis.md", "text": "https://wpnews.pro/news/dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis.txt", "jsonld": "https://wpnews.pro/news/dllm-tts-block-discrete-diffusion-language-model-for-text-to-speech-synthesis.jsonld"}}