10x Cheaper TTS at 50ms Time-to-First-Audio Nari Labs launched a Free Public Beta of realtime text-to-speech endpoints built on its own inference engine for the Qwen3-TTS model, claiming 50 ms time to first audio in Fast mode versus 300 ms for ElevenLabs TTS v3 and a 10x cost advantage in Standard mode at $5 per 1M characters against ElevenLabs V3's $50 per 1M characters. The free endpoints, qwen3-tts-fast:free and qwen3-tts:free, each allow up to 100 requests per day with a model-agnostic concurrency limit of 2, and Nari Labs said input streaming over websocket and custom voice cloning are coming soon. Nari Labs said it built the inference engine from scratch because existing engines were not fast or cheap enough for the model, and it is offering Early Access for higher concurrency and input streaming ahead of general availability. Since the beginning of 2024, open-weight LLMs have brought a new wind into the AI market with cheaper, yet competitive models. Now, trillions of tokens are being served using models from Z.ai, Kimi, Qwen, and DeepSeek. We believe that moment is coming to multimodal AI as well, starting with speech. Earlier this year, we realized the bottleneck was not the model, it was the inference engine. Qwen3-TTS is a phenomenal model when used correctly, but running it on existing engines is not fast and cheap enough. That’s why we built one from scratch blog post /blog/qwen3-tts-speed-cost-frontier/ , specifically tuned for the model. Our realtime TTS endpoints expand the Pareto frontier: - The world’s fastest TTS endpoint , with 50 ms time to first audio in Fast mode compared with 300 ms for ElevenLabs TTS v3.