cd /news/artificial-intelligence/narrate-complete-books-using-qwen3-t… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-98600] src=github.com β†— pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Narrate Complete Books Using Qwen3 TTS, Natural Sounding

A new open-source Rust project, tts-rs, delivers local text-to-speech on Apple silicon with three voice-cloning engines ported from PyTorch, achieving real-time factors as low as 0.252 for the Qwen3TTS engine on an M4/16 GB Mac. The project's benchmark shows a 1612-word chapter narrated in 2 minutes 40 seconds (11:36 audio) using Qwen3TTS, versus 5m 47s for Audio8 and 8m 15s for CosyVoice, with all engines running without Python at runtime. The author, drmhse, highlights that a 16-hour book would take about 4 hours of compute on a laptop, and notes a known but undiagnosed performance regression in Audio8's codec and CosyVoice's vocoder.

read6 min views6 publishedAug 16, 2026
Narrate Complete Books Using Qwen3 TTS, Natural Sounding
Image: Michielbdejong (auto-discovered)

Local text-to-speech in Rust. Text in, speech out, one process, no Python at runtime β€” three voice-cloning speech models ported from PyTorch and validated stage by stage against fp32 activations dumped from their references.

Requirements: macOS on Apple silicon. The custom kernels are Metal, and every number here was measured on an M4 / 16 GB. It builds and runs correctly elsewhere with --no-default-features

(CPU fallbacks, unit-tested), but that path is a portability guarantee rather than a deployment target β€” Audio8 measures RTF 2.151 on CPU against 0.554 on Metal, and it is the codec that suffers most (1.235 of it, against 0.214 on Metal).

./scripts/bootstrap.sh          # toolchain, all three checkpoints, assets, build

cargo run -p tts-cli --release -- speak \
    --engine audio8 --voice voices/cosy-default \
    --text "Hello from a fresh checkout." --out hello.wav

One command, nothing manual. bootstrap.sh

downloads and converts the checkpoints, fetches the fixtures for every gate, and builds. Name the engines you want β€” all three is ~13 GB, one is ~4 GB:

./scripts/bootstrap.sh --list              # the ids, their models, what each costs
./scripts/bootstrap.sh qwen3tts            # just one
./scripts/bootstrap.sh audio8 cosyvoice    # two

Every step is skipped if its output already exists, so re-running is cheap. Details in ** docs/reference.md**.

Clone a voice and speak | a 10-second reference clip becomes a tracked voice asset; no encoder at runtime | Serve an HTTP API | tts-serve β€” one engine loaded once, 3.0 s start, per-request voice and seed, cost headers on every response | Use it as a library | one Engine trait, engines chosen by string id at request time | Narrate a whole book | markdown in, delivery audio plus word-level timings out, resumable per stage |

All three narrating the same 1612-word chapter (examples/chapter.txt

, in the repo), each in the configuration it ships in, in two cloned voices, on one M4 / 16 GB laptop with no Python running. GitHub cannot embed audio in markdown β€” the demo page plays all six in place:

engine RTF wall time audio produced reach for it when
audio8
0.527–0.536 5m 47s 11:34 / 10:59 you want 44.1 kHz β€” the highest-fidelity output here
cosyvoice
0.703–0.718 8m 15s 12:48 / 11:44 you want the widest language coverage
qwen3tts
0.252–0.261
2m 40s
11:36 / 10:34 you are narrating something long

That bottom row is the point of the project: a chapter becomes 11 minutes of speech in 3 minutes, on a laptop. A 16-hour book is about 4 hours of compute rather than 12.

Compare the wall-time column, not just RTF. The three do not produce the same duration from the same text β€” cosyvoice

speaks slowest, 12:48 against audio8

's 11:34 β€” and since RTF divides by audio produced, a slower-speaking engine flatters its own RTF.

cargo run -p tts-cli --release -- speak --engine qwen3tts \
    --voice voices/cosy-default-qwen3tts --quant f16 \
    --text-file examples/chapter.txt --out chapter.wav

qwen3tts

gets there by batching across sections, which needs --quant f16

and needs length β€” on a 7-segment passage it is the slowest of the three at 0.665. It also supports ten languages only (en, de, es, zh, ja, fr, ko, ru, it, pt). The other two do not batch meaningfully and are steady at any length; audio8

is 2.36Γ— its PyTorch reference like-for-like, that reference running on MPS too.

Short-passage figures, for comparison β€” examples/senior.txt

, 132 words, median of five with the engines interleaved: audio8

0.554, cosyvoice

0.726, qwen3tts

0.665.

Ten languages on, a closed list: en, de, es, zh, ja, fr, ko, ru, it, pt. Text outside it has no faithful path through that engine.qwen3tts

A known performance regression, undiagnosed.audio8

's codec andcosyvoice

's vocoder are 35% and 32% slower than when first measured, while every transformer stage is unchanged. Both are convolution-heavy; the cause is likely the channels-last conv path.Sampled output is not reproducible across implementations, because the reference draws from torch's RNG. Pass aseed

for repeatability within this port; use the greedy path if you need to compare against PyTorch.

One comparison worth not making: cosyvoice

looks 6Γ— faster than its PyTorch reference, but only because upstream hardcodes cuda if available else cpu

and has no MPS path, so CPU is all it can do. Against a service that does use MPS it is ahead by about 5%.

TTS_API_KEY=secret cargo run -p tts-serve --release -- --port 3003

curl -X POST localhost:3003/tts -H "X-API-Key: secret" \
     -H 'content-type: application/json' \
     -d '{"text":"Hello from Rust.","voice":"voices/cosy-default-male","seed":7}' \
     -o out.wav -D headers.txt
route
POST /tts
WAV body, PCM s16le mono
POST /tts/stream
same, buffered rather than incremental
GET /v1/capabilities
engines, sample rates, and the weight formats each supports
GET /health
liveness
GET /
lists the live routes and the unimplemented ones, which answer 501

Every response carries its own cost. x-audio-seconds

, x-wall-seconds

, x-rtf

, and x-stages

with the per-stage split (llm=10.296,flow=25.129,vocoder=2.898

), so a client sees where the time went without a second request. ** voice and seed are per request** β€” the first selects a voice asset without a restart, the second makes a render reproducible.

scripts/narrate-book.sh --book path/to/document --out narration --engine qwen3tts
scripts/verify-narration.py narration/*.webm

One engine load for the whole run, resumable per stage (a section with a WAV master is never re-synthesised), deterministic under a seed. A 16-hour document is about 4 hours of synthesis at qwen3tts

's 0.260, against ~12 at cosyvoice

's 0.726, plus an hour of recognition either way.

use tts_core::{EngineConfig, SynthesisRequest, Voice};

let config = EngineConfig::new(tts_engines::default_root("cosyvoice"));
let engine = tts_engines::load("cosyvoice", &config)?;

let voice = Voice::load("voices/cosy-default-cosyvoice")?;
let request = SynthesisRequest::new("Hello from Rust.").with_voice(voice);
engine.validate(&request)?;                 // rejects a mismatched asset up front

let out = engine.synthesize(&request)?;
tts_core::wav::write("hello.wav", &out.audio)?;
println!("RTF {:.3}", out.stats.rtf(out.audio.seconds()));

Everything else is one file: ** docs/reference.md**.

|

ArchitectureValidationPerformancePorting trapsServing and narrationWhat did not work

crates/tts-core/        the Engine trait, voice assets, segmentation, WAV, the PRNG
crates/tts-nn/          shared model machinery + the custom Metal kernels
crates/tts-engines/     the registry β€” the one place that knows which engines exist
crates/tts-cli/         the `tts` binary: engines / voice / speak
crates/tts-serve/       the HTTP service: one engine, loaded once, behind a semaphore
crates/tts-bench/       the thermally-honest measurement harness
crates/tts-probe/       op-level benchmarks, one binary per question
crates/{audio8,cosyvoice,qwen3tts}/   one engine each, plus its fixture gate

references/{audio8,cosyvoice,qwen3tts}/   the PyTorch side: conversion, fixtures, quality
fixtures/{audio8,cosyvoice,qwen3tts}/     per-stage ground truth the gates compare against
voices/                 voice assets, one directory each (tracked β€” they are small)
examples/               senior.txt and chapter.txt, the two benchmark fixtures
scripts/                bootstrap, fetch-assets, gates, render-examples, narration

Everything shared is named tts-*

; everything engine-specific is named for its engine, and crates/audio8

, crates/cosyvoice

and crates/qwen3tts

match the ids --engine

takes.

Weights and virtualenvs are not tracked and docs/reference.md builds them. Fixtures are fetched rather than regenerated β€” scripts/fetch-assets.sh

pulls ~130 MB of checksummed ground truth from drmhse/tts-rs-assets, so

./scripts/gates.sh

runs without any PyTorch.Apache-2.0 (LICENSE). Contains no model weights β€” you download those yourself, and all three are Apache-2.0 per their model cards: Audio8-TTS-Preview-0.6b, Fun-CosyVoice3-0.5B, Qwen3-TTS-12Hz-1.7B-Base.

Two kinds of derived artifact are distributed: the voice assets in voices/

, encoded from reference audio those models ship, and the fixtures in drmhse/tts-rs-assets. The upstream licences reach both; attribution and the statement of changes are in

NOTICE.

── more in #artificial-intelligence 4 stories Β· sorted by recency
── more on @tts-rs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/narrate-complete-boo…] indexed:0 read:6min 2026-08-16 Β· β€”