{"slug": "stepaudio-3-gen-technical-report", "title": "StepAudio 3 Gen Technical Report", "summary": "Researchers submitted the StepAudio 3 Gen technical report to arXiv on 11 Sep 2026, introducing a general-purpose audio generation model that supports zero-shot text-to-speech, voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types in a unified framework. StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm, with a StepAudio Tokenizer representing general audio at 12.5 Hz in a shared 16 x 2048 residual code space. The report states the model achieves state-of-the-art performance on both TTS and voice design while retaining strong generation capabilities across speech, vocals, sound effects, and music.", "body_md": "# Computer Science > Sound\n\n  [Submitted on 11 Sep 2026]\n\n# Title:StepAudio 3 Gen Technical Report\n\n[View PDF](https://arxiv.org/pdf/2609.12945)\n\n[HTML (experimental)](https://arxiv.org/html/2609.12945v1)\n\nAbstract:We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \\times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at [this https URL](https://stepaudiollm.github.io/step-audio-3-gen/).\n    \n\n### Current browse context:\n\ncs.SD\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/stepaudio-3-gen-technical-report", "canonical_source": "https://arxiv.org/abs/2609.12945", "published_at": "2026-09-15 15:35:39+00:00", "updated_at": "2026-09-15 15:50:41.507131+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "natural-language-processing", "ai-research"], "entities": ["StepAudio 3 Gen", "StepAudio Tokenizer", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/stepaudio-3-gen-technical-report", "markdown": "https://wpnews.pro/news/stepaudio-3-gen-technical-report.md", "text": "https://wpnews.pro/news/stepaudio-3-gen-technical-report.txt", "jsonld": "https://wpnews.pro/news/stepaudio-3-gen-technical-report.jsonld"}}