StepAudio 3 Gen Technical Report Researchers submitted the StepAudio 3 Gen technical report to arXiv on 11 Sep 2026, introducing a general-purpose audio generation model that supports zero-shot text-to-speech, voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types in a unified framework. StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm, with a StepAudio Tokenizer representing general audio at 12.5 Hz in a shared 16 x 2048 residual code space. The report states the model achieves state-of-the-art performance on both TTS and voice design while retaining strong generation capabilities across speech, vocals, sound effects, and music. Computer Science Sound Submitted on 11 Sep 2026 Title:StepAudio 3 Gen Technical Report View PDF https://arxiv.org/pdf/2609.12945 HTML experimental https://arxiv.org/html/2609.12945v1 Abstract:We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech TTS , voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization RVQ tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: 1 interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, 2 RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and 3 discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at this https URL https://stepaudiollm.github.io/step-audio-3-gen/ . Current browse context: cs.SD References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .