Narrate Complete Books Using Qwen3 TTS, Natural Sounding A new open-source Rust project, tts-rs, delivers local text-to-speech on Apple silicon with three voice-cloning engines ported from PyTorch, achieving real-time factors as low as 0.252 for the Qwen3TTS engine on an M4/16 GB Mac. The project's benchmark shows a 1612-word chapter narrated in 2 minutes 40 seconds (11:36 audio) using Qwen3TTS, versus 5m 47s for Audio8 and 8m 15s for CosyVoice, with all engines running without Python at runtime. The author, drmhse, highlights that a 16-hour book would take about 4 hours of compute on a laptop, and notes a known but undiagnosed performance regression in Audio8's codec and CosyVoice's vocoder. Local text-to-speech in Rust. Text in, speech out, one process, no Python at runtime — three voice-cloning speech models ported from PyTorch and validated stage by stage against fp32 activations dumped from their references. Requirements: macOS on Apple silicon. The custom kernels are Metal, and every number here was measured on an M4 / 16 GB. It builds and runs correctly elsewhere with --no-default-features CPU fallbacks, unit-tested , but that path is a portability guarantee rather than a deployment target — Audio8 measures RTF 2.151 on CPU against 0.554 on Metal, and it is the codec that suffers most 1.235 of it, against 0.214 on Metal . ./scripts/bootstrap.sh toolchain, all three checkpoints, assets, build cargo run -p tts-cli --release -- speak \ --engine audio8 --voice voices/cosy-default \ --text "Hello from a fresh checkout." --out hello.wav One command, nothing manual. bootstrap.sh downloads and converts the checkpoints, fetches the fixtures for every gate, and builds. Name the engines you want — all three is ~13 GB, one is ~4 GB: ./scripts/bootstrap.sh --list the ids, their models, what each costs ./scripts/bootstrap.sh qwen3tts just one ./scripts/bootstrap.sh audio8 cosyvoice two Every step is skipped if its output already exists, so re-running is cheap. Details in docs/reference.md . Clone a voice and speak | a 10-second reference clip becomes a tracked voice asset; no encoder at runtime | Serve an HTTP API | tts-serve — one engine loaded once, 3.0 s start, per-request voice and seed, cost headers on every response | Use it as a library | one Engine trait, engines chosen by string id at request time | Narrate a whole book | markdown in, delivery audio plus word-level timings out, resumable per stage | All three narrating the same 1612-word chapter examples/chapter.txt , in the repo , each in the configuration it ships in, in two cloned voices, on one M4 / 16 GB laptop with no Python running. GitHub cannot embed audio in markdown — the demo page https://drmhse.github.io/tts-rs/ plays all six in place: | engine | RTF | wall time | audio produced | reach for it when | |---|---|---|---|---| audio8 | 0.527–0.536 | 5m 47s | 11:34 / 10:59 | you want 44.1 kHz — the highest-fidelity output here | cosyvoice | 0.703–0.718 | 8m 15s | 12:48 / 11:44 | you want the widest language coverage | qwen3tts | 0.252–0.261 | 2m 40s | 11:36 / 10:34 | you are narrating something long | That bottom row is the point of the project: a chapter becomes 11 minutes of speech in 3 minutes, on a laptop. A 16-hour book is about 4 hours of compute rather than 12. Compare the wall-time column, not just RTF. The three do not produce the same duration from the same text — cosyvoice speaks slowest, 12:48 against audio8 's 11:34 — and since RTF divides by audio produced, a slower-speaking engine flatters its own RTF. cargo run -p tts-cli --release -- speak --engine qwen3tts \ --voice voices/cosy-default-qwen3tts --quant f16 \ --text-file examples/chapter.txt --out chapter.wav qwen3tts gets there by batching across sections, which needs --quant f16 and needs length — on a 7-segment passage it is the slowest of the three at 0.665. It also supports ten languages only en, de, es, zh, ja, fr, ko, ru, it, pt . The other two do not batch meaningfully and are steady at any length; audio8 is 2.36× its PyTorch reference like-for-like, that reference running on MPS too. Short-passage figures, for comparison — examples/senior.txt , 132 words, median of five with the engines interleaved: audio8 0.554, cosyvoice 0.726, qwen3tts 0.665. Ten languages on , a closed list: en, de, es, zh, ja, fr, ko, ru, it, pt. Text outside it has no faithful path through that engine. qwen3tts A known performance regression, undiagnosed. audio8 's codec and cosyvoice 's vocoder are 35% and 32% slower than when first measured, while every transformer stage is unchanged. Both are convolution-heavy; the cause is likely the channels-last conv path. Sampled output is not reproducible across implementations , because the reference draws from torch's RNG. Pass a seed for repeatability within this port; use the greedy path if you need to compare against PyTorch. One comparison worth not making: cosyvoice looks 6× faster than its PyTorch reference, but only because upstream hardcodes cuda if available else cpu and has no MPS path, so CPU is all it can do. Against a service that does use MPS it is ahead by about 5%. TTS API KEY=secret cargo run -p tts-serve --release -- --port 3003 curl -X POST localhost:3003/tts -H "X-API-Key: secret" \ -H 'content-type: application/json' \ -d '{"text":"Hello from Rust.","voice":"voices/cosy-default-male","seed":7}' \ -o out.wav -D headers.txt | route | | |---|---| POST /tts | WAV body, PCM s16le mono | POST /tts/stream | same, buffered rather than incremental | GET /v1/capabilities | engines, sample rates, and the weight formats each supports | GET /health | liveness | GET / | lists the live routes and the unimplemented ones, which answer 501 | Every response carries its own cost. x-audio-seconds , x-wall-seconds , x-rtf , and x-stages with the per-stage split llm=10.296,flow=25.129,vocoder=2.898 , so a client sees where the time went without a second request. voice and seed are per request — the first selects a voice asset without a restart, the second makes a render reproducible. scripts/narrate-book.sh --book path/to/document --out narration --engine qwen3tts scripts/verify-narration.py narration/ .webm One engine load for the whole run, resumable per stage a section with a WAV master is never re-synthesised , deterministic under a seed. A 16-hour document is about 4 hours of synthesis at qwen3tts 's 0.260, against ~12 at cosyvoice 's 0.726, plus an hour of recognition either way. js use tts core::{EngineConfig, SynthesisRequest, Voice}; let config = EngineConfig::new tts engines::default root "cosyvoice" ; let engine = tts engines::load "cosyvoice", &config ?; let voice = Voice::load "voices/cosy-default-cosyvoice" ?; let request = SynthesisRequest::new "Hello from Rust." .with voice voice ; engine.validate &request ?; // rejects a mismatched asset up front let out = engine.synthesize &request ?; tts core::wav::write "hello.wav", &out.audio ?; println "RTF {:.3}", out.stats.rtf out.audio.seconds ; Everything else is one file: docs/reference.md . | Architecture /drmhse/tts-rs/blob/main/docs/reference.md architecture Validation /drmhse/tts-rs/blob/main/docs/reference.md validation Performance /drmhse/tts-rs/blob/main/docs/reference.md performance Porting traps /drmhse/tts-rs/blob/main/docs/reference.md porting-traps Serving and narration /drmhse/tts-rs/blob/main/docs/reference.md serving-and-narration What did not work /drmhse/tts-rs/blob/main/docs/reference.md what-did-not-work crates/tts-core/ the Engine trait, voice assets, segmentation, WAV, the PRNG crates/tts-nn/ shared model machinery + the custom Metal kernels crates/tts-engines/ the registry — the one place that knows which engines exist crates/tts-cli/ the tts binary: engines / voice / speak crates/tts-serve/ the HTTP service: one engine, loaded once, behind a semaphore crates/tts-bench/ the thermally-honest measurement harness crates/tts-probe/ op-level benchmarks, one binary per question crates/{audio8,cosyvoice,qwen3tts}/ one engine each, plus its fixture gate references/{audio8,cosyvoice,qwen3tts}/ the PyTorch side: conversion, fixtures, quality fixtures/{audio8,cosyvoice,qwen3tts}/ per-stage ground truth the gates compare against voices/ voice assets, one directory each tracked — they are small examples/ senior.txt and chapter.txt, the two benchmark fixtures scripts/ bootstrap, fetch-assets, gates, render-examples, narration Everything shared is named tts- ; everything engine-specific is named for its engine, and crates/audio8 , crates/cosyvoice and crates/qwen3tts match the ids --engine takes. Weights and virtualenvs are not tracked and docs/reference.md /drmhse/tts-rs/blob/main/docs/reference.md setup builds them. Fixtures are fetched rather than regenerated — scripts/fetch-assets.sh pulls ~130 MB of checksummed ground truth from drmhse/tts-rs-assets https://huggingface.co/datasets/drmhse/tts-rs-assets , so ./scripts/gates.sh runs without any PyTorch.Apache-2.0 LICENSE /drmhse/tts-rs/blob/main/LICENSE . Contains no model weights — you download those yourself, and all three are Apache-2.0 per their model cards: Audio8-TTS-Preview-0.6b https://huggingface.co/Audio8/Audio8-TTS-Preview-0.6b , Fun-CosyVoice3-0.5B https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B , Qwen3-TTS-12Hz-1.7B-Base https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base . Two kinds of derived artifact are distributed: the voice assets in voices/ , encoded from reference audio those models ship, and the fixtures in drmhse/tts-rs-assets https://huggingface.co/datasets/drmhse/tts-rs-assets . The upstream licences reach both; attribution and the statement of changes are in NOTICE /drmhse/tts-rs/blob/main/NOTICE .