# CrispASR: One C++ binary for multilingual ASR and TTS models

> Source: <https://github.com/CrispStrobe/CrispASR>
> Published: 2026-08-31 22:22:23+00:00

**One C++ binary, 54 ASR backends + 55 TTS engines + multilingual text translation, zero Python dependencies.**

CrispASR started as a fork of [whisper.cpp](https://github.com/ggml-org/whisper.cpp) and extends that base into a **unified speech engine** called `crispasr`

, backed by full ggml C++ runtimes for major open-weights ASR *and* TTS architectures. One build, one binary, one consistent CLI — pick the backend at the command line or let CrispASR auto-detect it from your GGUF file. See [Text-to-Speech](#text-to-speech-models) for the TTS side.

``` bash
$ crispasr -m ggml-base.en.bin          -f samples/jfk.wav                    # OpenAI Whisper
$ crispasr -m parakeet-tdt-0.6b.gguf    -f samples/jfk.wav                    # NVIDIA Parakeet
$ crispasr -m canary-1b-v2.gguf         -f samples/jfk.wav                    # NVIDIA Canary
$ crispasr -m voxtral-mini-3b-2507.gguf -f samples/jfk.wav                    # Mistral Voxtral
$ crispasr --backend qwen3 -m auto      -f samples/jfk.wav                    # -m auto downloads
$ crispasr --backend kokoro -m auto --tts "Hello world" --tts-output out.wav  # TTS
```

No Python. No PyTorch. No separate per-model binary. No `pip install`

. Just one C++ binary and a GGUF file.

**Browser**: All backends compile to WebAssembly (4.3 MB) via `build-wasm.sh`

.
Multithreaded, runs entirely client-side with COOP/COEP headers.

**Demo**: [HuggingFace Space](https://huggingface.co/spaces/cstr/CrispASR) —
live transcription + TTS + language detection, auto-deployed from `hf-space/`

.

| Project | What it does |
|---|---|
|
This repo — C++ speech engine. 54 ASR + 55 TTS backends, CLI + HTTP server + C-ABI + Python/Rust/Dart/Go/Ruby/Java bindings. |
|
Cross-platform Flutter transcription app built on CrispASR. Desktop + mobile, model browser with download queue, mic capture, SRT/VTT/JSON export, diarization, batch processing. Fully offline. |
|
Text-related engine via ggml — same philosophy as CrispASR but for embeddings, retrieval, OCR and OMR, Math and Music Notation. Numerous architectures (XLM-R, Qwen3-Embed, Gemma3, ModernBERT, ...), dense + sparse + ColBERT + reranking. PP-OCR, Tesseract, EasyOCR, InternVL2, etc. Python/Rust/Dart bindings. |
|
Python ASR GUI with 9 backends (faster-whisper, mlx-whisper, voxtral, insanely-fast-whisper, ...). The Python counterpart to CrispASR's C++ approach. |

— new to CrispASR? Two commands to your first working audio, no repo clone needed**Start here**[Supported backends](#supported-backends)—[ASR](#asr-backends)+[TTS](#text-to-speech-models)+[translation](#translation)+[post-processing](#post-processing-models)+[music & audio analysis](#music--audio-analysis)[Feature matrix](#feature-matrix)[Install & build](#install--build)— quick install (full guide in[docs/install.md](/CrispStrobe/CrispASR/blob/main/docs/install.md));— the[which prebuilt Linux tarball to download](/CrispStrobe/CrispASR/blob/main/docs/install.md#prebuilt-linux-tarballs--which-one-to-download-355)`-hip`

/`-vulkan`

builds require the matching GPU driver and do**not** fall back to CPU (the`-cuda`

tarballs do, since v0.8.30)[Quick start — ASR](#quick-start)[Troubleshooting](/CrispStrobe/CrispASR/blob/main/docs/troubleshooting.md)— it printed the banner and stopped, reading the exit code,`--no-gpu`

bisect, which Windows zip— 52 engines: Kokoro, Qwen3-TTS, VibeVoice, dots.tts, Orpheus, Chatterbox, IndexTTS, Irodori, VoxCPM2, CosyVoice3, CSM, Dia, Zonos, Bark, Piper, MeloTTS, and more**Text-to-Speech (TTS)**[Streaming & live transcription](/CrispStrobe/CrispASR/blob/main/docs/streaming.md)[Server mode (HTTP API)](/CrispStrobe/CrispASR/blob/main/docs/server.md)[Concurrency, parallelism & scaling](/CrispStrobe/CrispASR/blob/main/docs/concurrency.md)— how one transcription uses multiple cores, concurrent server requests (`--server-workers`

), bulk offline transcription, replicas behind a load balancer[CLI reference](/CrispStrobe/CrispASR/blob/main/docs/cli.md)— flags, VAD, CTC alignment, output formats, auto-download, audio formats[Environment variables](/CrispStrobe/CrispASR/blob/main/docs/environment-variables.md)— the`CRISPASR_<BACKEND>_<FEATURE>`

convention, global knobs, and every per-backend variable[Language bindings](/CrispStrobe/CrispASR/blob/main/docs/bindings.md)— Python / Rust / Dart / Go / Java / JavaScript / Ruby / mobile[Benchmarking CrispASR](/CrispStrobe/CrispASR/blob/main/docs/benchmarking.md)— how to measure transcribe time (not cold start): server/in-process reps, proof-of-work rules, phase-timing env vars[Architecture](/CrispStrobe/CrispASR/blob/main/docs/architecture.md)— layered layout,`src/core/`

primitives, regression discipline[Contributing — adding a new backend](/CrispStrobe/CrispASR/blob/main/docs/contributing.md)— 5-file recipe, ground-truth diff workflow[Regression matrix](/CrispStrobe/CrispASR/blob/main/docs/regression-matrix.md)—`tools/test-all-backends.py`

capability tiers— synthetic-audio marking (watermark + C2PA + spoken disclaimer), what counts as a voice clone, the speaker-biometrics boundary, why there is no emotion recognition, and what stays your duty as deployer**EU AI Act**[Quantize models](/CrispStrobe/CrispASR/blob/main/docs/quantize.md)—`crispasr-quantize`

for all backends[GPU backend selection](#gpu-backend-selection)[Debugging & profiling](#debugging--profiling)[Credits](#credits)

Everything below this section is a catalogue — 100+ backends, browse it when you
need one. If you just want CrispASR *working*, this is the whole path. No repo
clone, no Python, no model hunting.

Download one file from [ Releases](https://github.com/CrispStrobe/CrispASR/releases/latest)
and unzip it:

| Platform | Download | Notes |
|---|---|---|
Windows |
`crispasr-windows-x86_64-cpu.zip` |
Needs AVX2 (2013+ Intel / 2015+ AMD). Older CPU → `…-cpu-legacy.zip` |
Windows + NVIDIA |
`crispasr-windows-x86_64-cuda.zip` |
Self-contained; a CUDA Toolkit install is not required. CUDA-13-native build: `…-cuda13.zip` (Turing+) |
macOS |
`crispasr-macos.tar.gz` |
Metal GPU support built in |
Linux |
`crispasr-linux-x86_64.tar.gz` |
`…-cuda.tar.gz` / `…-vulkan.tar.gz` for GPU |

Prefer to build it yourself? See [Install & build](#install--build). The `-hip`

and `-vulkan`

builds require the matching driver and do **not** fall back to CPU;
the Linux `-cuda`

tarballs do fall back.

Check it runs — this should print a version banner and exit:

```
crispasr --version          # Windows: .\crispasr.exe --version
```

`-m auto`

downloads the model on first use (~135 MB here) and reuses it
afterwards — nothing to find or install. It lands in `~/.cache/crispasr/`

(`%USERPROFILE%\.cache\crispasr`

on Windows).

```
crispasr --backend kokoro -m auto --tts "The quick brown fox jumps over the lazy dog." --tts-output hello.wav
# crispasr: TTS output written to 'hello.wav' (78000 samples @ 24000 Hz, 3.25 sec)
```

Play `hello.wav`

. That is the TTS half working.

```
crispasr --backend parakeet -m auto -f hello.wav -l en
# crispasr: transcribed 3.2s audio in 0.32s (10.1x realtime)
# The quick brown fox jumps over the lazy dog.
```

(~467 MB on first run. `-l en`

skips language auto-detection, which would
otherwise fetch a small extra model.) Both halves now work — swap in your own
`.wav`

and you are running.

| You want to… | Go to |
|---|---|
| Clone a voice from a recording |
`--i-have-rights` |

[Which backend should I pick?](#which-backend-should-i-pick)[docs/cli.md](/CrispStrobe/CrispASR/blob/main/docs/cli.md)[docs/streaming.md](/CrispStrobe/CrispASR/blob/main/docs/streaming.md)[docs/server.md](/CrispStrobe/CrispASR/blob/main/docs/server.md)`crispasr --list-backends`

Add `-v`

to any command for verbose progress, and `--dry-run-resolve`

to print
which model files it would open (and whether they're on disk) without loading
anything.

If a command printed its banner and then simply stopped — no error, no output
file — that is a crash, not a refusal, and the exit code identifies it in one
step. See ** docs/troubleshooting.md**.

CrispASR ships **54 ASR backends** for transcription/translation and
**55 TTS engines** for synthesis. It also ships audio-to-audio S2S backends,
including Sidon restoration and the VoxCPM2 AudioVAE speech upscaler; see the [feature matrix](/CrispStrobe/CrispASR/blob/main/docs/feature-matrix.md)
for the complete capability list.
Pick at the CLI with `--backend NAME`

, or omit it to let the binary auto-detect
from the GGUF metadata. Jump to the [TTS table](#text-to-speech-models) for the synthesis side.

| Backend | Model | Architecture | Languages | License |
|---|---|---|---|---|
whisper |
`ggml-base.en.bin` |

**whisper**`distil-whisper/distil-large-v3`

**parakeet**`nvidia/parakeet-tdt-0.6b-v3`

**parakeet**`nvidia/parakeet-tdt-0.6b-v2`

**parakeet**`nvidia/parakeet-tdt-1.1b`

**parakeet**`nvidia/parakeet-tdt_ctc-110m`

**parakeet**`nvidia/parakeet-tdt_ctc-1.1b`

**parakeet**`nvidia/parakeet-tdt_ctc-0.6b-ja`

**reazonspeech**`reazon-research/reazonspeech-nemo-v2`

**fastconformer-ctc**`nvidia/parakeet-ctc-0.6b`

**fastconformer-ctc**`nvidia/parakeet-ctc-1.1b`

**fastconformer-ctc**`grider-transwithai/parakeet-ctc-1.1b-ja`

**canary**`nvidia/canary-1b-v2`

`-sl/-tl`

)**canary-qwen**`nvidia/canary-qwen-2.5b`

**lfm2-audio**`LiquidAI/LFM2.5-Audio-1.5B`

**lfm2-audio**`LiquidAI/LFM2.5-Audio-1.5B-JP`

**mini-omni2**`gpt-omni/mini-omni2`

**cohere**`CohereLabs/cohere-transcribe-03-2026`

**cohere**`efwkjn/cohere-asr-ja-v0.1`

**granite**[,](https://huggingface.co/ibm-granite/granite-speech-3.3-2b)`ibm-granite/granite-speech-{3.2-8b,3.3-2b,3.3-8b}`

`granite-4.0-1b-speech`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#granite--granite-41--granite-41-plus--granite-41-nar))**granite-4.1**`ibm-granite/granite-speech-4.1-2b`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#granite--granite-41--granite-41-plus--granite-41-nar))**granite-4.1-plus**`ibm-granite/granite-speech-4.1-2b-plus`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#granite--granite-41--granite-41-plus--granite-41-nar))**granite-4.1-nar**`ibm-granite/granite-speech-4.1-2b-nar`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#granite--granite-41--granite-41-plus--granite-41-nar))**fastconformer-ctc**`nvidia/stt_en_fastconformer_ctc_large`

**voxtral**`mistralai/Voxtral-Mini-3B-2507`

**voxtral4b**`mistralai/Voxtral-Mini-4B-Realtime-2602`

**qwen3**`Qwen/Qwen3-ASR-0.6B`

**qwen3-1.7b**`Qwen/Qwen3-ASR-1.7B`

**qwen3-ja-anime**`jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame-hf`

**mega-asr**`zhifeixie/Mega-ASR`

**higgs-stt**`bosonai/higgs-audio-v3-stt`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#higgs-stt))**wav2vec2**`jonatasgrosman/wav2vec2-large-xlsr-53-english`

**wav2vec2**`facebook/data2vec-audio-base-960h`

**wav2vec2**`facebook/hubert-large-ls960-ft`

**glm-asr**`zai-org/GLM-ASR-Nano-2512`

**kyutai-stt**`kyutai/stt-1b-en_fr`

**kyutai-stt**`kyutai/stt-2.6b-en`

**firered-asr**`FireRedTeam/FireRedASR2-AED`

**moonshine**`UsefulSensors/moonshine-{tiny,base}`

**moonshine‑de**`fidoriel/moonshine-base-de`

**moonshine‑tiny‑de**`fidoriel/moonshine-tiny-de`

**moonshine-streaming**`UsefulSensors/moonshine-streaming-{tiny,small,medium}`

**gemma4-e2b**`google/gemma-4-E2B-it`

**gemma4-e4b**`google/gemma-4-E4B-it`

`--backend gemma4-e2b`

**omniasr**`omniASR-CTC-1B-v2`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#omniasr-ctc--llm--unlimited))**1600+****omniasr‑300m**`omniASR-CTC-300M-v2`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#omniasr-ctc--llm--unlimited))**1600+****omniasr-llm**`omniASR-LLM-300M-v2`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#omniasr-ctc--llm--unlimited))**1600+****omniasr-llm**`omniASR-LLM-Unlimited-300M-v2`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#omniasr-ctc--llm--unlimited))**1600+****vibevoice**`microsoft/VibeVoice-ASR`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#vibevoice))**vibevoice-bitnet**`VibeVoice-ASR-BitNet`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#vibevoice))**mimo-asr**`XiaomiMiMo/MiMo-V2.5-ASR`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#mimo-asr))**ark-asr*** experimental/WIP*[(base](https://huggingface.co/cstr/ark-asr-3b-GGUF)`cstr/ark-asr-3b-GGUF`

[)](https://huggingface.co/AutoArk-AI/ARK-ASR-3B)`AutoArk-AI/ARK-ASR-3B`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#ark-asr))**moss-audio**`OpenMOSS-Team/MOSS-Audio-4B-Instruct`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#moss-audio))**moss-transcribe**`OpenMOSS-Team/MOSS-Transcribe-preview-2B`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#moss-transcribe))**moss-diarize**`OpenMOSS-Team/MOSS-Transcribe-Diarize-0.9B`

**whisper***(tiron)** experimental*[(base](https://huggingface.co/cstr/tiron-GGML)`Trelis/tiron`

[)](https://huggingface.co/Trelis/tiron)`Trelis/tiron`

**funasr**`FunAudioLLM/Fun-ASR-Nano-2512`

**fun-asr-mlt-nano**`FunAudioLLM/Fun-ASR-MLT-Nano-2512`

**paraformer**`funasr/paraformer-zh`

**foxnose***(speaker diarization)*`Wespeaker/wespeaker-voxceleb-resnet34-LM`

`--diarize-method foxnose`

: WeSpeaker ResNet34-LM 256-d embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi temporal smoothing ([more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#foxnose-diarize)). 3.18 % DER on VoxConverse dev vs the upstream reference implementation's 3.07 %**gigaam**[(base](https://huggingface.co/cstr/gigaam-v3-GGUF)`ai-sage/GigaAM-v3`

[)](https://huggingface.co/ai-sage/GigaAM-v3)`ai-sage/GigaAM-v3`

`e2e_rnnt`

/ `e2e_ctc`

emit punctuation + casing + ITN from a SentencePiece vocab, `rnnt`

/ `ctc`

emit bare lowercase Cyrillic ([more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#gigaam))**sensevoice**`FunAudioLLM/SenseVoiceSmall`

`-oj`

JSON expose the tags as separate fields. Upstream's emotion classifier is **not exposed**— see[EU AI Act](/CrispStrobe/CrispASR/blob/main/docs/eu-ai-act.md#41-emotion-recognition--removed-not-gated)| Backend | Model | Architecture | Input / output | License |
|---|---|---|---|---|
sidon |
`KevinAHM/Sidon-GGUF` |

[)](https://huggingface.co/sarulab-speech/sidon-v0.1)

`sarulab-speech/sidon-v0.1`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#sidon))**voxcpm2-vae**[, converted with](https://huggingface.co/openbmb/VoxCPM2)`openbmb/VoxCPM2`

`--vae-only`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#voxcpm2-vae))

```
huggingface-cli download KevinAHM/Sidon-GGUF sidon-v0.1-f16.gguf --local-dir models
crispasr -m models/sidon-v0.1-f16.gguf -f input.wav --s2s --s2s-output restored.wav

python models/convert-voxcpm2-to-gguf.py --input openbmb/VoxCPM2 \
  --output models/voxcpm2-vae-f32.gguf --vae-only
crispasr -m models/voxcpm2-vae-f32.gguf -f input.wav --s2s \
  --s2s-output upscaled.wav
```

Synthesis backends, driven by the `--tts`

flag and a `--tts-output PATH.wav`

.
See the dedicated [Text-to-Speech](#text-to-speech-models) section below for
quick-start commands and engine selection guidance.

| Backend | Models | Architecture | Languages | License |
|---|---|---|---|---|
miotts |
`MioTTS-0.6B` |

**vibevoice-tts**[,](https://huggingface.co/cstr/vibevoice-realtime-0.5b-GGUF)`VibeVoice-Realtime-0.5B`

`VibeVoice-1.5B`

**kugelaudio**`kugelaudio-0-open`

**qwen3-tts**[,](https://huggingface.co/cstr/qwen3-tts-0.6b-base-GGUF)`Qwen3-TTS-12Hz-0.6B-Base`

[,](https://huggingface.co/cstr/qwen3-tts-1.7b-base-GGUF)`1.7B-Base`

`1.7B-VoiceDesign`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#qwen3-tts))**qwen3-tts-customvoice**`1.7B-CustomVoice`

`--voice <name>`

); optional style via `--instruct`

(e.g. "spoke very slowly") ([more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#qwen3-tts))**moss-tts**`OpenMOSS-Team/MOSS-TTS-v1.5`

`--voice ref.wav`

; `--backend moss-tts -m <backbone> --codec-model <codec>`

**moss-tts-local**`OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5`

`--backend moss-tts-local -m <backbone> --codec-model <codec>`

**omnivoice**`k2-fsa/OmniVoice`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#omnivoice))**melotts**[EN_V2](https://github.com/myshell-ai/MeloTTS)`myshell-ai/MeloTTS`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#melotts))**piper**[community voices](https://github.com/rhasspy/piper)`rhasspy/piper`

`--g2p-dict`

)**kokoro**[+ German backbones](https://huggingface.co/hexgrad/Kokoro-82M)`hexgrad/Kokoro-82M`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#kokoro))**orpheus**[+](https://huggingface.co/cstr/orpheus-3b-0.1-ft-GGUF)`Orpheus-3B-FT`

`SNAC 24 kHz`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#orpheus))**chatterbox**[+ Nano/turbo/fine-tune variants](https://huggingface.co/cstr/chatterbox-GGUF)`cstr/chatterbox-GGUF`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#chatterbox--chatterbox-turbo--chatterbox-nano--chatterbox-finnish-nano--kartoffelbox-turbo--lahgtna-chatterbox))`chatterbox-finnish-nano`

) fine-tunes**indextts**`cstr/indextts-1.5-GGUF`

**voxcpm2-tts**`cstr/voxcpm2-GGUF`

`--voice <wav>`

**voxtral-tts**`mistralai/Voxtral-4B-TTS-2603`

**cosyvoice3-tts**`cstr/cosyvoice3-0.5b-2512-GGUF`

`--voice <name>`

, or any WAV via `--voice ref.wav --ref-text "<exact transcript>"`

. `--backend cosyvoice3-tts-rl`

selects upstream's RL-tuned talker (same companions)**csm**`cstr/csm-1b-GGUF`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#csm))**lfm2-audio**[+](https://huggingface.co/cstr/lfm2-audio-1.5b-GGUF)`cstr/lfm2-audio-1.5b-GGUF`

`jp`

**dia**`nari-labs/Dia-1.6B`

`[S1]`

/`[S2]`

tags (use >100-char prompts)**zonos-tts**[+](https://huggingface.co/cstr/zonos-v0.1-transformer-GGUF)`cstr/zonos-v0.1-transformer-GGUF`

`cstr/dac-44khz-GGUF`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#zonos-tts))**bark**`cstr/bark-small-GGUF`

`.npz`

prompts (`--voice <file.npz>`

)**speecht5**`cstr/speecht5-tts-GGUF`

`--voice <xvector.bin>`

)**fastpitch**`cstr/fastpitch-en-GGUF`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#fastpitch))**bananamind-tts**`Banaxi-Tech/BananaMind-TTS-V2.1-Preview`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#bananamind-tts))**parler-tts**`cstr/parler-tts-mini-v1.1-GGUF`

`--instruct`

)**outetts**`cstr/outetts-0.3-1b-GGUF`

`--voice <speaker.json>`

)**pocket-tts**`cstr/pocket-tts-GGUF`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#pocket-tts))**tada**[+](https://huggingface.co/cstr/tada-tts-1b-GGUF)`cstr/tada-tts-1b-GGUF`

`HumeAI/tada-3b-ml`

`tada-ref.gguf`

, custom voices via `--voice <tada-ref.gguf>`

built with `models/convert-tada-ref-to-gguf.py`

([more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#tada))**TTS feature matrix**

| Backend | Voice cloning | Sampling | kHz | Auto-download | Flash attn |
|---|---|---|---|---|---|
| vibevoice-tts | yes | temp | 24 | yes | yes |
| qwen3-tts | yes* | temp | 24 | yes | yes |
| omnivoice | yes | temp | 24 | — | — |
| kokoro | — | — | 24 | yes | — |
| orpheus | — | temp | 24 | yes | yes |
| chatterbox | yes | temp | 24 | yes | yes |
| outetts | yes (JSON) | temp | 24 | yes | yes |
| indextts | yes | temp | 24 | yes | yes |
| voxcpm2-tts | yes | — | 48 | yes | — |
| cosyvoice3-tts | yes | temp | 24 | yes | yes |
| f5-tts | yes | — | 24 | yes | — |
| irodori-tts | yes (WAV) | VoiceDesign: `--instruct` |
48 | yes | — |
| csm | — | temp | 24 | yes | — |
| dia | — | temp | 44 | yes | — |
| bark | yes (.npz) | temp | 24 | yes | — |
| speecht5 | yes (xvec) | — | 16 | yes | — |
| parler-tts | — | temp | 44 | yes | — |
| fastpitch | — | — | 22 | — | — |
| piper | — | — | 22 | — | — |
| pocket-tts | yes | temp | 24 | yes | — |
| tada | yes | temp | 24 | yes | — |
| dots-tts | yes (`--voice ref.wav` ) |
16-step CFG Euler | 48 | yes | — |
| confucius4-tts | yes (`--voice ref.wav` ) |
25-step CFG Euler | 22.05 | yes | — |

* CustomVoice variant only; Base uses baked speakers via `--voice <name>`

.

**Output language.** `-tl <lang>`

(or `-l`

) selects the language to speak;
`cosyvoice3-tts`

, `qwen3-tts`

and `moss-tts`

act on it natively. For
cross-lingual **cloning** — an English reference clip speaking German, the
subtitle-dubbing case — also pass `-sl <lang>`

for the language the reference is
spoken in, so cosyvoice3 drops the reference transcript instead of carrying its
accent. Over HTTP: `"language"`

+ `"source_lang"`

on `POST /v1/audio/speech`

.
See [ docs/tts.md](/CrispStrobe/CrispASR/blob/main/docs/tts.md#output-language-and-cross-lingual-cloning--tl---sl).

Text-to-text translation, distinct from the audio-side `--translate`

flag (which routes audio → English text on whisper / canary / etc.).
Driven by `--text "..." -sl <src> -tl <tgt>`

.

| Backend | Models | Architecture | Languages | License |
|---|---|---|---|---|
m2m100 |
`facebook/m2m100_418M` |

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#m2m100--wmt21))**m2m100-wmt21**[+](https://huggingface.co/cstr/wmt21-dense-24-wide-en-x-GGUF)`facebook/wmt21-dense-24-wide-en-x`

`facebook/wmt21-dense-24-wide-x-en`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#m2m100--wmt21))`en-x`

/ `x-en`

checkpoints)**madlad**`google/madlad400-3b-mt`

[more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#madlad))

```
# m2m100 base (production-ready)
./build/bin/crispasr --backend m2m100 -m auto \
    --text "Hello world, how are you today?" \
    -sl en -tl de
# → Hallo Welt, wie bist du heute?

# WMT21 dense (English ↔ X, 4.7B — auto-downloads ~2.5 GB).
# Two separate checkpoints: en-x for English-source, x-en for
# English-target. Pick the one matching your `-sl`/`-tl` direction
# (or pass an explicit `-m <path>` to load the other manually).
./build/bin/crispasr --backend m2m100-wmt21 -m auto \
    --text "The president said he would not attend." \
    -sl en -tl de   # uses wmt21-dense-24-wide-en-x

./build/bin/crispasr --backend m2m100-wmt21 \
    -m models/wmt21-dense-24-wide-x-en-q4_k.gguf \
    --text "Le président a dit qu'il ne serait pas présent." \
    -sl fr -tl en   # uses wmt21-dense-24-wide-x-en

# MADLAD-400 3B (419 languages, bit-token-identical to Python SP)
./build/bin/crispasr --backend madlad -m auto \
    --text "Hello world." \
    -sl en -tl ta
```

For 2-stage pipelines (e.g., ASR → m2m100), use the dedicated
`--tr-sl`

/ `--tr-tl`

flags; they fall back to `-sl`

/ `-tl`

when
unset, so single-stage standalone usage is just `-sl/-tl`

.

Work with all backends.

| Model | Task | Architecture | Languages | License | HuggingFace |
|---|---|---|---|---|---|
FireRedPunc |
Punctuation restoration | BERT-base (12L, d=768), 5 classes | Chinese + English | Apache-2.0 |
`cstr/fireredpunc-GGUF` |

**fullstop-punc**`cstr/fullstop-punc-multilang-GGUF`

**punctuate-all**`cstr/punctuate-all-GGUF`

**PCS**`--punc-model pcs`

**truecaser‑lstm**`--truecase-model lstm`

**truecaser‑crf**`--truecase-model crf`

**truecaser‑de**`--truecase-model auto`

**CLD3**`cstr/cld3-GGUF`

**GlotLID-V3**`cstr/glotlid-GGUF`

**LID-176**`cstr/fasttext-lid176-GGUF`

Shared codec modules used by TTS backends. Also available standalone for encode/decode.

| Model | Architecture | Sample Rate | Token Rate | License | HuggingFace |
|---|---|---|---|---|---|
MioCodec v2 |
WavLM encoder → FSQ(12800) → Transformer decoder + AdaLN-Zero + SnakeBeta upsampler + iSTFT | 44.1 kHz | 25 Hz (341 bps) | MIT |
`cstr/miocodec-v2-44k-GGUF` |

**SNAC 24 kHz**`cstr/snac-24khz-GGUF`

All runtimes share ggml-based inference. The speech-LLM backends (**qwen3**, **voxtral**, **voxtral4b**, **granite**, **glm-asr**, **kyutai-stt**) inject audio encoder frames directly into an autoregressive language model's input embeddings, instead of using a dedicated CTC/transducer/seq2seq decoder. The **fastconformer-ctc** backend hosts the NeMo FastConformer-CTC standalone ASR family — `stt_en_fastconformer_ctc_{large,xlarge,xxlarge}`

and the architecturally-identical `parakeet-ctc-{0.6b,1.1b}`

(different training data + tokenizer, same encoder + head shape) — with greedy CTC decoding. Same C++ runtime as the canary-ctc aligner.

Beyond speech, CrispASR runs several music/audio analysis tasks — each a small
GGUF with the architecture auto-detected, no Python. See [ docs/cli.md](/CrispStrobe/CrispASR/blob/main/docs/cli.md)
for the per-task flags and output formats.

**Source separation**(`--separate`

) — split a mix into stems (`<input>_<stem>.wav`

) via**mel-band-roformer**(vocal/instrumental, MIT) or** htdemucs**(4-stem).`--stems vocals,drums`

selects a subset;`--sep-output-dir`

sets the output location.**Piano transcription**(`--backend piano-transcription`

) — piano audio → MIDI note events (88 keys @ 100 fps, ByteDance/Kong CRNN; F16 GGUF ≈ 77 MB).**Guitar tablature**(`--tab`

) — per-frame fret-per-string grid via**TabCNN**(Wiggins & Kim, ISMIR 2019; CC BY 4.0 weights). The backend emits per-string emission scores, not a decided tablature — run your own constrained Viterbi via`crispasr_session_tab_emissions()`

for playable output.**Beat / downbeat tracking**(`--beats`

) — beat grid via**Beat This!**(CPJKU, ISMIR 2024; MIT for code*and*weights, no patent-encumbered DBN).**Chord recognition**(`--chords`

) — chord timeline (`.lab`

) via**BTC**(ISMIR 2019). Weights are CC-BY-NC-SA, gated behind`--accept-license cc-by-nc-sa-4.0`

.**Pitch / F0 estimation**(`--pitch`

) — monophonic pitch track via**CREPE**(MIT).

Run `crispasr --list-backends`

to see it live. Each backend declares capabilities at runtime; if you ask for a feature the selected backend does not support, CrispASR prints a warning and silently ignores the flag.

**Sortable / filterable view:** [ docs/feature-matrix.html](/CrispStrobe/CrispASR/blob/main/docs/feature-matrix.html) — click any column header to sort, type to filter rows, click cap pills to require a capability. Generated from

`crispasr --list-backends-json`

(single source of truth — drift impossible). Regenerate via `python tools/gen-feature-matrix.py`

. A Markdown twin lives at [.](/CrispStrobe/CrispASR/blob/main/docs/feature-matrix.md)

`docs/feature-matrix.md`

The static table below is a curated subset focusing on the ASR backends and the cross-cutting features that matter for ASR pipelines. The full 109-backend × 27-cap surface is in the generated views.

| Feature | whisper | parakeet | canary | cohere | granite | granite‑4.1 | voxtral | voxtral4b | qwen3 | fc‑ctc | wav2vec2 | glm‑asr | kyutai‑stt | firered | moonshine | moon‑stream | omniasr | omniasr‑llm | vibevoice | gemma4‑e2b | mimo‑asr | funasr | paraformer | sensevoice |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Native timestamps | ✔ | ✔ | ✔ | ✔ | ✔ | |||||||||||||||||||
| CTC timestamps | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ||||
| Word-level timing | ✔ | ✔ | ✔ | ✔ | `-am` |
✔† | `-am` |
`-am` |
`-am` |
`-am` |
`-am` |
`-am` |
✔ | `-am` |
`-am` |
`-am` |
`-am` |
`-am` |
`-am` |
`-am` |
`-am` |
`-am` |
||
| Per-token confidence | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ||||
| Language auto-detect | ✔ | ✔ | LID | LID | LID | LID | LID | LID | ✔ | LID | LID | ✔ | LID | LID | LID | LID | LID | LID | LID | ✔ | LID | LID | LID | ✔ |
| Speech translation | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ||||||||||||||||||
| Speaker diarization | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Grammar (GBNF) | ✔ | |||||||||||||||||||||||
| Temperature sampling | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | |||||
| Beam search | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | |||||||||||||
| Flash attention | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ||
| Punctuation toggle | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | |||||||||
| Punc restoration | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp | pp |
| Source / target language | ✔ | ✔ | ✔ | ✔ | ✔ | |||||||||||||||||||
Audio Q&A (`--ask` ) |
* | * | ✔ | * | * | * | * | |||||||||||||||||
| Streaming | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
Auto-download (`-m auto` ) |
✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
KV quant (`CRISPASR_KV_QUANT` , plus per-half `_K` / `_V` ) |
✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ||||||||||||||
mmap weights (`CRISPASR_GGUF_MMAP` ) |
✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | |
| TTS | ✔ |

The matrix above covers 24 ASR backends. **Additional ASR backends** not shown: `nemotron`

(39-lang streaming ASR with cache-aware FastConformer + RNN-T), `lfm2-audio`

(ASR + TTS + S2S in one model), `moss-audio`

(audio understanding + ASR), `moss-transcribe`

(Qwen3-Omni encoder + Qwen3-1.7B ASR), `mini-omni2`

(ASR + TTS + S2S), `kugelaudio`

(7B audio understanding). See [ docs/feature-matrix.md](/CrispStrobe/CrispASR/blob/main/docs/feature-matrix.md) for the full 109-backend matrix.

**TTS-only backends**(

`kokoro`

, `qwen3-tts`

+ variants, `vibevoice-tts`

, `orpheus`

+ DE variants, `chatterbox`

/ `chatterbox-turbo`

/ `chatterbox-nano`

/ `kartoffelbox-turbo`

/ `lahgtna-chatterbox`

, `dia`

, `bark`

, `outetts`

, `zonos`

, `csm`

, `f5-tts`

, `irodori-tts`

, `parler-tts`

, `speecht5`

, `piper`

, `fastpitch`

, `pocket-tts`

, `melotts`

, `cosyvoice3`

, `voxcpm2`

, `tada-tts`

) all carry the TTS, AUTO_DOWNLOAD, TEMPERATURE, and FLASH_ATTN caps; per-backend cloning + voice-pack support is documented in the [Text-to-Speech models](#text-to-speech-models)table above and

[. The vibevoice and lfm2-audio columns mark dual-mode (ASR + TTS) backends.](/CrispStrobe/CrispASR/blob/main/docs/tts.md)

`docs/tts.md`

**Key:** ✔ = native/built-in, `-am`

= via CTC forced aligner (`-am canary-ctc-aligner.gguf`

or `-am qwen3-forced-aligner.gguf`

), **LID** = via external language identification pre-step (`-l auto`

), **pp** = via `--punc-model`

post-processor (FireRedPunc or fullstop-punc), * = experimental or partial support, † = PLUS variant only (native `[T:N]`

word timestamps with `-owts`

; base uses `-am`

). granite-4.1 covers both the regular and `-plus`

variants; granite-4.1-nar is a non-autoregressive variant with encoder+projector only (no LLM decode features). The **KV quant** row marks backends that honor `CRISPASR_KV_QUANT={f16,q8_0,q4_0}`

— CTC-style backends without a KV cache (parakeet, fc-ctc, wav2vec2, kyutai-stt, firered, moonshine variants, omniasr-CTC) don't apply. The same backends also honor the per-half `CRISPASR_KV_QUANT_K`

/ `CRISPASR_KV_QUANT_V`

overrides (llama.cpp `--cache-type-k`

/ `--cache-type-v`

parity) for asymmetric K-vs-V precision; common recipe `K=q8_0 V=q4_0`

saves ~40 % more KV memory than symmetric Q8_0. The **mmap weights** row marks backends consuming `core_gguf::load_weights()`

and therefore honoring `CRISPASR_GGUF_MMAP=1`

; whisper itself uses upstream's loader and is unaffected. See [ docs/cli.md](/CrispStrobe/CrispASR/blob/main/docs/cli.md) Memory footprint for usage + recommended combos.

**Speaker diarization** as a post-processing step via `--diarize`

:

`energy`

/`xcorr`

— stereo-only, no extra deps`foxnose`

—**best accuracy, no external deps**: WeSpeaker ResNet34-LM embeddings + GMM/BIC speaker counting + spectral clustering + Viterbi smoothing. Estimates the speaker count rather than needing it up front;`--diarize-embedder auto`

fetches the GGUF (24 MB, CC-BY-4.0). 7.3 % DER on VoxConverse dev where`pyannote`

+ TitaNet scores 7.8 %, and 3.18 % vs the upstream reference's 3.07 % when scored on turns ([more](/CrispStrobe/CrispASR/blob/main/docs/architecture.md#foxnose-diarize))`pyannote`

— native GGUF (no Python, no sherpa-onnx); add`--diarize-embedder auto`

(TitaNet) or`--diarize-embedder indextts`

(ECAPA-TDNN) for globally stable speaker IDs across long files`sherpa`

/`ecapa`

— external[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)subprocess; runs once globally on full audio for consistent speaker IDs (#110)`vad-turns`

— mono-friendly gap-based proxy

The server endpoint supports `response_format=diarized_json`

for structured speaker-labelled output with normalised speaker letters (A, B, C …) — see [ docs/server.md](/CrispStrobe/CrispASR/blob/main/docs/server.md#diarized-json-format-206).

Full reference + tuning knobs (cluster threshold, max speakers, pluggable embedder adapters): see [ docs/cli.md#diarization](/CrispStrobe/CrispASR/blob/main/docs/cli.md#diarization).

**Language identification** for backends without native LID: `--lid-backend whisper`

(default, 75 MB ggml-tiny.bin), `--lid-backend silero`

(native GGUF, 16 MB, 95 languages), or `--lid-backend firered`

(FireRedLID, 1.7 GB, 120 languages — Conformer encoder + Transformer decoder).

**Voice activity detection**: `--vad`

uses the default Silero VAD (~885 KB, auto-downloaded). Each VAD segment is transcribed independently, producing separate SRT/VTT entries with correct timestamps. Use `--vad --split-on-punct`

for best subtitle output. Four VAD backends: Silero (default), FireRedVAD (`-vm firered`

, recommended), MarbleNet (`-vm marblenet`

, 439 KB, 6 languages), Whisper-VAD-EncDec (`-vm whisper-vad`

, experimental).

**Punctuation restoration** (`--punc-model`

): CTC-based backends output lowercase without punctuation. Named shortcuts: `auto`

/`firered`

(Chinese+English), `fullstop`

(EN/DE/FR/IT, XLM-R-large), `punctuate-all`

(12 languages, XLM-R-base), `pcs`

(47 languages, punc + truecasing + sentence boundary detection in one model). Or pass a GGUF path directly. Also available via Python/Rust/Dart wrappers (`crispasr.PuncModel`

).

**Truecasing** (`--truecase-model`

): Restore German noun/name capitalization in lowercase ASR output. Three options in ascending quality: `auto`

(statistical, 1.7 MB), `crf`

(CRF with context, 8.5 MB), `lstm`

(BiLSTM char-level, 3.2 MB, **recommended** — 97.9% F1, handles adjective/noun distinction and formal "Ihnen"). All auto-download from [ cstr/truecaser-de](https://huggingface.co/cstr/truecaser-de). Or use

`--punc-model pcs`

for neural punc + truecasing in one pass (47 languages).## Which backends produce punctuation natively?

| Backend | Punctuation | Capitalization | Notes |
|---|---|---|---|
| whisper | ✔ | ✔ | Full punctuation and casing |
| parakeet | ✔ | ✔ | |
| canary | ✔ | ✔ | |
| cohere | ✔ | ✔ | Toggleable via `--no-punctuation` |
| granite | ✔ | ✔ | LLM output |
| voxtral | ✔ | ✔ | LLM output |
| voxtral4b | ✔ | ✔ | LLM output |
| qwen3 | ✔ | ✔ | LLM output |
| funasr | ✔ | ✔ | LLM output (Qwen3-0.6B decoder). Chinese chars carry full-width period; mlt-nano variant adds Latin-script casing + punctuation. |
| sensevoice | ✔ | ✔ | CTC output with native ITN — toggle off via `--no-punctuation` , which controls Arabic-digit vs spelled-out numerals + comma/period emission. |
| paraformer | no |
no |
NAR character-level output — add `--punc-model` |
| gigaam | ✔ (`e2e_*` ) |
✔ (`e2e_*` ) |
The `e2e_*` revisions carry punctuation + casing + inverse text normalization in the SentencePiece vocabulary. The charwise `ctc` / `rnnt` revisions emit lowercase Cyrillic with no punctuation — but auto-restoration is still suppressed for them, because the auto-enabled FireRedPunc is a Chinese/English model and injects full-width CJK punctuation into Russian. Use an `e2e_*` revision for punctuated output, or pass an explicit `--punc-model` . |
| glm-asr | ✔ | ✔ | LLM output |
| kyutai-stt | ✔ | ✔ | LLM output |
| moonshine | ✔ | ✔ | Encoder-decoder output |
fastconformer-ctc |
no |
no |
CTC — add `--punc-model` |
wav2vec2 |
no |
no |
CTC — add `--punc-model` |
firered-asr |
no |
no |
CTC — add `--punc-model` |
omniasr (CTC) |
no |
no |
CTC — add `--punc-model` |
omniasr (LLM) |
✔ | ✔ | Autoregressive decoder |

Other freely-licensed alternatives that could be added: [felflare/bert-restore-punctuation](https://huggingface.co/felflare/bert-restore-punctuation) (MIT, English, includes truecasing), [xashru/punctuation-restoration](https://github.com/xashru/punctuation-restoration) (Apache-2.0, 40+ languages, BiLSTM-CRF).

**Progressive subtitle output** (`--flush-after`

): By default, non-whisper backends buffer all segments and print output at the end. For real-time subtitle consumption (PotPlayer, custom media players), use `--flush-after 1`

to print each SRT entry to stdout immediately after its VAD segment is transcribed:

```
crispasr --backend parakeet -m parakeet.gguf --vad --flush-after 1 -osrt -f long_audio.wav
# SRT entries appear progressively as each segment finishes
```

**JSON output with language detection**: When using `-l auto -oj`

, the JSON output includes detected language info:

```
{
  "crispasr": {
    "backend": "cohere",
    "language": "en",
    "language_detected": "en",
    "language_confidence": 0.977,
    "language_source": "ecapa"
  },
  "transcription": [...]
}
```

| Need | Pick |
|---|---|
| Battle-tested, all features exposed | whisper |
| Lowest English WER | cohere |
Fastest (16x realtime on CPU) |
moonshine (tiny), fc-ctc (10x) |
| Multilingual + word timestamps + fast | parakeet (2.9x RT) |
Multilingual with explicit language control |
canary |
Speech translation (X→en or en→X) |
canary, voxtral, qwen3 |
30 languages + Chinese dialects |
qwen3 |
1600+ languages |
omniasr (CTC or LLM) |
Realtime streaming ASR (native incremental encoder, ~2× RT feed; sub-second-token target deferred to phase 2) |
voxtral4b |
| Highest-quality offline speech-LLM | voxtral |
| Apache-licensed speech-LLM | granite, voxtral, qwen3, omniasr-llm |
Lightweight CTC-only (fast, no decoder) |
wav2vec2, fc-ctc, data2vec, omniasr |
Russian |
gigaam (`e2e_rnnt` — 8.4 % avg WER, punctuation + ITN), whisper, qwen3 |
Mandarin + Chinese dialects |
firered-asr, qwen3, glm-asr, funasr, paraformer, sensevoice |
Multilingual (31 langs) speech-LLM |
fun-asr-mlt-nano, qwen3, omniasr-llm, gemma4-e2b |
Multilingual (50+ langs) + LID + audio-event in one pass |
sensevoice (encoder-only CTC, non-AR, 15× faster than Whisper-Large) |

Audio-LLM backends (`qwen3`

, `voxtral`

, `granite`

, `glm-asr`

, etc.) run full
transformer decoder stacks (28+ layers, 2048-dim) and are **dramatically slower
on CPU** than encoder-only backends. On older dual-core hardware they can drop
below 0.01× realtime. If you're on CPU-only hardware:

- Prefer
**moonshine**(16× RT),** fc-ctc**(10× RT),** parakeet**(2.9× RT), or** whisper**for usable speeds. - Use
`--flush-after 1`

to see results as each VAD slice completes instead of waiting for the entire file. - Use
`-pp`

/`--print-progress`

for per-slice progress indicators on all backends (unified backends show slice-level progress; whisper shows encoder-level progress). - Quantize models to Q4_K or Q5_K to reduce memory and compute.

Cohere, canary, granite, voxtral and voxtral4b need an explicit
language code up front. If you don't know the language, pass
`-l auto`

and crispasr runs an optional LID pre-step before the main
transcribe() call:

```
# Downloads ggml-tiny.bin (75 MB, 99 languages) on first use
crispasr --backend cohere -m $TC/cohere-transcribe-q5_0.gguf \
         -f unknown.wav -l auto
# crispasr[lid]: detected 'en' (p=0.977) via whisper-tiny
# crispasr: LID -> language = 'en' (whisper, p=0.977)
```

These LID providers are available:

-
`--lid-backend whisper`

(default) — uses a small multilingual ggml-*.bin model via the crispasr C API. Auto-downloads ~75 MB on first use. 99 languages. -
`--lid-backend silero`

— native GGUF port of Silero's 95-language classifier. 16 MB F32. Runs as a ggml graph (multi-threaded SIMD on CPU, GPU offload on Metal/CUDA; on Vulkan the graph is routed to CPU pending an upstream kernel fix). Analyzes the first 30 s of audio (`CRISPASR_SILERO_LID_MAX_S`

overrides);`CRISPASR_SILERO_LID_LEGACY=1`

restores the old scalar path. -
`--lid-backend ecapa`

—**recommended**: ECAPA-TDNN (Apache-2.0). Purpose-built for language ID. Very high accuracy on TTS benchmark. Two variants via`--lid-model`

:— VoxLingua107, 43 MB F16, 107 languages, ISO codes (en, de, ...).`cstr/ecapa-lid-107-GGUF`

**Default.**— CommonLanguage, 40 MB F16, 45 languages, full names (English, German, ...).`cstr/ecapa-lid-commonlanguage-GGUF`

-
`--lid-backend firered`

— FireRedLID (Conformer encoder + Transformer decoder). Q4_K (544 MB), 120 languages including Chinese dialects. Slower but covers more languages. -
`--lid-backend probe`

— no second model at all: ask the**ASR model itself**. Transcribes a 20 s clip once per language the model declares and keeps the best-scoring candidate (length × text-LID agreement × distinct-token ratio², the last term catching the repetitive output a wrong-language prompt produces). Currently implemented by**cohere**. The reason to prefer it is correctness, not just the saved download: an external detector knows 99 languages while Cohere Transcribe accepts 14 — and its Arabic finetune only`en`

/`ar`

— so external LID regularly returns a language the model was never trained on, and Cohere answers a wrong language*fluently*rather than failing. The probe cannot. Cost is one encode + one short decode per candidate, so it runs automatically only for a model with ≤ 4 languages (`CRISPASR_COHERE_PROBE_MAX_LANGS`

);`CRISPASR_COHERE_PROBE_TEXTLID=0`

drops the text-LID agreement term.**The ceiling is about cost, not accuracy.** Measured on the real models: the two-language Arabic finetune picks`ar`

for an Arabic clip (p=0.675) and`en`

for`samples/jfk.wav`

(p=0.647); the 14-language base model, probed across all 14, also gets both right (`en`

p=0.169,`ar`

p=0.254) — it is simply slower than an external detector. The encoder output is language-independent, so the probe encodes**once** and decodes per candidate (encode is ~87 % of a pass); a 14-candidate probe measured**12 s → 4-5 s** against one-encode-per-candidate, byte-identical output.`CRISPASR_COHERE_PROBE_REUSE_ENC=0`

restores the naive path.The one soft spot worth knowing: asking the model for a language it was

*not*trained on can yield a clean translation rather than garbage, which a text LID then confirms — "fluent French out" is not evidence of French in. That is what makes a mismatched whitelist dangerous, and it is measurable: force a 14-language list onto the two-language Arabic finetune and its`fr`

probe returns real French and wins. The real base model's`fr`

probe instead code-switches ("Et so, my fellow Americans…", agreement 0.00) and loses, as it should.

These VAD providers are available:

**Silero VAD**(default) — ~885 KB, auto-downloaded via`--vad`

. Industry-standard, well-tested.**FireRedVAD**— DFSMN-based, 2.4 MB, F1=97.57%. Pass`--vad -vm firered`

to auto-download. Recommended.**MarbleNet**— NVIDIA 1D separable CNN, 439 KB, 6 languages (EN/DE/FR/ES/RU/ZH). Pass`--vad -vm marblenet`

to auto-download. Smallest model. ()`cstr/marblenet-vad-GGUF`

**Whisper-VAD-EncDec***(experimental)*— Whisper-base encoder + TransformerDecoder head, 22 MB Q4_K. Trained on Japanese ASMR; may not generalise well to all domains. Pass`--vad -vm whisper-vad`

. Slower than others (~1s vs ~50ms). ()`cstr/whisper-vad-encdec-asmr-GGUF`

Pass `--lid-backend off`

to skip LID entirely.

Audio LID (above) tags **what was spoken**; text LID tags **what was
written**. Text LID runs on a transcript or any UTF-8 string and is
useful for routing post-ASR pipelines (translation, punctuation, sub
selection) without re-running an audio model. Three GGUF families,
one binary — the dispatcher picks by `general.architecture`

:

| Backend | Labels | Size (F16) | License | HF repo |
|---|---|---|---|---|
CLD3 (Google compact language detector v3) |
109 ISO 639-1 | 440 KB |
Apache-2.0 |
`cstr/cld3-GGUF` |

**GlotLID-V3**(cis-lmu fastText)`cstr/glotlid-GGUF`

**LID-176**(Facebook fastText)`cstr/fasttext-lid176-GGUF`

¹ LID-176 is **CC-BY-NC-4.0** — non-commercial use only. CLD3 +
GlotLID-V3 are Apache-2.0 with no such constraint. Pick CLD3 for the
smallest, fastest path; GlotLID for maximum coverage (low-resource
languages); LID-176 only if you need its specific 176-label space and
accept its non-commercial terms.

**Standalone CLI** — auto-routes by GGUF arch, with auto-download:

```
crispasr-lid -m auto --text "Bonjour le monde"        # → cstr/cld3-GGUF (default, ~440 KB)
crispasr-lid -m auto:glotlid --text "Bonjour le monde" -k 5
crispasr-lid -m auto:lid-fasttext176 --text "Hallo Welt"
# Or pass an explicit path / canonical filename (looked up in the registry):
crispasr-lid -m cld3-f16.gguf --text "你好世界"
# zh	0.997816
echo "Привет мир" | crispasr-lid -m auto --quiet
# ru	0.907322
```

**Post-ASR pipeline** — `--lid-on-transcript`

runs the same dispatcher
on the assembled transcript (also accepts `auto[:variant]`

):

```
crispasr -m ggml-tiny.bin -f speech.wav --lid-on-transcript auto
# (transcript on stdout)
# lang=de	conf=0.997123	backend=lid-cld3
```

The dispatcher (`src/text_lid_dispatch.{h,cpp}`

) is a thin C ABI
façade — one integer compare per call; per-stage diff harness is
green at cos≥0.999 across 8 multilingual smoke samples.

**Don't want to build?** Prebuilt binaries for Windows, macOS and Linux are on
the [releases page](https://github.com/CrispStrobe/CrispASR/releases/latest) —
see [Start here](#start-here) for which file to take. The rest of this section
is for building from source.

```
git clone --recursive https://github.com/CrispStrobe/CrispASR
cd CrispASR
# already cloned without --recursive? initialize the bundled ggml submodule:
#   git submodule update --init --recursive
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
```

The `ggml/`

submodule is required. If you cloned without `--recursive`

, run
`git submodule update --init --recursive`

first — otherwise CMake stops with a
message telling you to do exactly that.

Produces `build/bin/crispasr`

(main CLI), `build/bin/crispasr-quantize`

,
and `build/bin/crispasr-diff`

. No Python, PyTorch, or pip required at
runtime — just a C++17 compiler and CMake 3.14+.

For GPU acceleration, add the matching ggml flag at configure time:

```
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON     # NVIDIA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON    # Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON   # cross-vendor
```

**See ** for the full guide:
all GPU backends (CUDA / Metal / Vulkan / MUSA / SYCL), Windows
convenience scripts, ffmpeg ingestion, optional BLAS, glibc notes,
and the

`docs/install.md`

`scripts/dev-build.sh`

wrapper.If a build runs but the binary exits with no output, see
[ docs/troubleshooting.md](/CrispStrobe/CrispASR/blob/main/docs/troubleshooting.md).

Deeper ASR examples below. If this is your first run, use
[Start here](#start-here) instead. For TTS, the runnable guide is
[docs/tts.md](/CrispStrobe/CrispASR/blob/main/docs/tts.md) ([Text-to-Speech](#text-to-speech-models) below is
the model catalogue).

```
# Download a whisper model (same as upstream whisper.cpp)
./models/download-ggml-model.sh base.en

./build/bin/crispasr -m models/ggml-base.en.bin -f samples/jfk.wav
# [00:00:00.000 --> 00:00:07.940]   And so my fellow Americans ask not what your country can do for you
# [00:00:07.940 --> 00:00:10.760]   ask what you can do for your country.
# Grab the quantized model (~467 MB)
curl -L -o parakeet.gguf \
    https://huggingface.co/cstr/parakeet-tdt-0.6b-v3-GGUF/resolve/main/parakeet-tdt-0.6b-v3-q4_k.gguf

./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav
# Auto-detected backend 'parakeet' from GGUF metadata.
# And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.

# Word-level timestamps (one line per word)
./build/bin/crispasr -m parakeet.gguf -f samples/jfk.wav -ml 1
# Transcription (source == target)
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl de

# Translation (German speech → English text)
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -sl de -tl en

# ...or use the familiar crispasr flag:
./build/bin/crispasr --backend canary -m canary-1b-v2-q5_0.gguf -f audio.de.wav -l de --translate
# First run downloads ~2.5 GB to ~/.cache/crispasr/ via curl, then runs
./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav

# Subsequent runs use the cached file
./build/bin/crispasr --backend voxtral -m auto -f samples/jfk.wav -l en
# 0.6B (default, ~500 MB)
./build/bin/crispasr --backend qwen3 -m auto -f audio.zh.wav

# 1.7B (higher quality, ~1.3 GB) — supports both -hf and non-hf source models
./build/bin/crispasr --backend qwen3 -m qwen3-1.7b --auto-download -f audio.wav

# Japanese anime/galgame fine-tune (~1.3 GB)
./build/bin/crispasr --backend qwen3 -m qwen3-ja-anime --auto-download -f anime.wav
```

**Long audio:** the default is safe 30 s chunking. `--chunk-seconds 0`

decodes
the whole file in ONE pass (matches the reference model verbatim on multi-minute
clips, #218) — but the encoder's full attention is O(N²) in audio length, so
keep single-pass clips under ~10 minutes on 16 GB machines. For long-form use
prefer the plain `-q4_k`

/`-q8_0`

GGUFs over the `-imatrix`

variants (see the
model card).

```
./build/bin/crispasr --backend glm-asr -m auto -f audio.wav

# Long audio in one pass (up to 655 s — 30 s encoder windows, one LLM prompt,
# same layout as the HF/zai reference; matches it verbatim on the #218 clip):
./build/bin/crispasr --backend glm-asr -m auto --chunk-seconds 0 -f long.wav
```

Note: in single-pass mode the model (like the reference) skips leading
non-speech audio; the default 30 s-chunked mode transcribes more of such
clips. Custom `--ask`

/ non-English `--language`

instructions need a GGUF
with baked BPE merges (re-published 2026-07; older GGUFs fall back to the
default transcription prompt with a warning).

```
# Download the LM + audio tokenizer (the tokenizer is a separate model)
huggingface-cli download cstr/mimo-asr-GGUF mimo-asr-q4_k.gguf \
    --local-dir ~/.cache/crispasr
huggingface-cli download cstr/mimo-tokenizer-GGUF mimo-tokenizer-q4_k.gguf \
    --local-dir ~/.cache/crispasr

# Transcribe (auto-discovers tokenizer if it sits next to the LM)
./build/bin/crispasr \
    --backend mimo-asr \
    -m ~/.cache/crispasr/mimo-asr-q4_k.gguf \
    --codec-model ~/.cache/crispasr/mimo-tokenizer-q4_k.gguf \
    -f samples/jfk.wav
# Output: And so, my fellow Americans, ask not what your country can do
# for you. Ask what you can do for your country.
```

The 4.5 GB Q4_K is the recommended quant; F16 (14.9 GB) needs ~16 GB
RAM during inference. JFK matches the upstream Python
`MimoAudio.asr_sft`

reference verbatim; performance on M1+Metal is
~0.3× realtime (Q4_K dequant per step is the bottleneck — F16 +
KV-reuse follow-ups are queued under PLAN #51a/b/c).

```
# English (Q4_K quantized, 212 MB — 6x smaller than F16)
curl -L -o wav2vec2-en-q4k.gguf \
    https://huggingface.co/cstr/wav2vec2-large-xlsr-53-english-GGUF/resolve/main/wav2vec2-xlsr-en-q4_k.gguf

./build/bin/crispasr -m wav2vec2-en-q4k.gguf -f samples/jfk.wav
# and so my fellow americans ask not what your country can do for you ask what you can do for your country

# German
curl -L -o wav2vec2-de-q4k.gguf \
    https://huggingface.co/cstr/wav2vec2-large-xlsr-53-german-GGUF/resolve/main/wav2vec2-xlsr-de-q4_k.gguf

./build/bin/crispasr -m wav2vec2-de-q4k.gguf -f audio.de.wav

# Convert any HuggingFace Wav2Vec2ForCTC model:
python models/convert-wav2vec2-to-gguf.py \
    --model-dir jonatasgrosman/wav2vec2-large-xlsr-53-german \
    --output wav2vec2-de.gguf --dtype f32
# Then optionally quantize:
./build/bin/crispasr-quantize wav2vec2-de.gguf wav2vec2-de-q4k.gguf q4_k
```

CrispASR has three feature areas that warrant their own docs pages:

—[Streaming & live transcription](/CrispStrobe/CrispASR/blob/main/docs/streaming.md)`--stream`

,`--mic`

,`--live`

, sliding-window chunking, per-token confidence.— Kokoro (multilingual, smallest), Qwen3-TTS (highest fidelity, voice cloning), VibeVoice (lowest-latency streaming), Orpheus (3 B Llama + SNAC), Chatterbox (flow-matching + HiFT vocoder, German via Kartoffelbox), IndexTTS, VoxCPM2, and CosyVoice3 (9 langs + 18 zh dialects; baked-voice bank + arbitrary-WAV cloning). Voice packs, language routing, and qwen3-tts environment switches. All TTS output is watermarked; post-embed verification warns if confidence is low. Use[Text-to-Speech (TTS)](/CrispStrobe/CrispASR/blob/main/docs/tts.md)`--detect-watermark file.wav`

to check any WAV for AI watermarks.— persistent model, OpenAI-compatible[Server mode (HTTP API)](/CrispStrobe/CrispASR/blob/main/docs/server.md)`/v1/audio/transcriptions`

(ASR) and`/v1/audio/speech`

+`/v1/voices`

(TTS, automatic on any loaded CAP_TTS backend), per-request voice + speed + instructions, CORS, long-form sentence chunking, API keys, Docker Compose, prebuilt CUDA images.— one transcription already uses multiple cores; the server accepts requests concurrently but serializes inference on one model by default;[Concurrency, parallelism & scaling](/CrispStrobe/CrispASR/blob/main/docs/concurrency.md)`--server-workers N`

runs N model instances so pure-ASR requests run concurrently; and for bulk/throughput workloads, process-level fan-out (`xargs -P`

/ GNU`parallel`

) or N replicas behind a load balancer. Also covers what is*not*supported (batched multi-stream inference, PagedAttention) and why.

Quickest taste of each:

```
# Streaming from microphone
crispasr --mic -m model.gguf

# TTS via auto-downloaded VibeVoice (~636 MB on first run)
crispasr --backend vibevoice-tts -m auto --tts "Hello world" --tts-output hello.wav

# CosyVoice3 on GPU; companions auto-download beside the LLM
crispasr --backend cosyvoice3-tts -m auto --tts "Hello world" --tts-output cosy.wav

# CosyVoice3 fast mode: 5 flow steps instead of the quality-default 10
COSYVOICE3_FLOW_STEPS=5 crispasr --backend cosyvoice3-tts -m auto \
  --tts "Hello world" --tts-output cosy-fast.wav

# Persistent HTTP server, OpenAI-compatible
crispasr --server -m model.gguf --port 8080
curl -F "file=@audio.wav" http://localhost:8080/v1/audio/transcriptions

# TTS over HTTP — load a TTS backend, hit /v1/audio/speech
crispasr --server --backend qwen3-tts-customvoice -m auto --voice-dir ./voices --port 8080
curl http://localhost:8080/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"input":"Hello world","voice":"vivian"}' -o out.wav
```

CosyVoice3 uses batched classifier-free guidance and request-sized KV
caching by default. Baked voices load only the LLM, flow, HiFT, and voice
bank; the larger S3 tokenizer and CAMPPlus companions load lazily when a
`.wav`

cloning voice is first requested.

Common flags:

```
crispasr -m auto --backend parakeet -f audio.wav --vad -osrt --split-on-punct
```

| Flag | Meaning |
|---|---|
`-m FNAME` / `--backend NAME` |
Model path (or `auto` ) and forced backend |
`-f FNAME` |
Input audio (repeatable; positional accepted) |
`--vad` |
Silero VAD chunking — strongly recommended for multi-minute audio |
`-osrt` / `-ovtt` / `-otxt` / `-oj` / `-ojf` |
Output formats (also `-ocsv` , `-olrc` ) |
`-am FNAME` |
CTC aligner GGUF for word-level timestamps on LLM backends |
`--align-only` |
Standalone forced alignment: text/`.srt` + audio → timestamped SRT/JSON (no ASR needed); `.srt` input keeps its cues and gets re-timed (`--align-granularity auto|word|segment` ) |
`-tp F` / `-bs N` |
Sampling temperature / beam search width |
`-n N` / `--frequency-penalty F` |
Generated-token cap / opt-in repeated-token penalty for supported autoregressive ASR backends |
`-l auto` / `--detect-language` |
LID pre-step for backends without native lang detect |
`--hotwords "A,B,C"` |
Contextual biasing — boost named terms during CTC/TDT decode or LLM prompt |
`-ck N` |
Fallback chunk size when VAD is off (default 30 s) |
`--list-backends` |
Print the capability matrix and exit |

**See ** for the full reference: every
flag, VAD details, CTC alignment workflow, output JSON layout, the
auto-download registry, and supported audio formats.

`docs/cli.md`

**See** for Python / Rust / Dart / Go / Java / JavaScript / Ruby / mobile.

`docs/bindings.md`

CrispASR is structured as a stable C-ABI in `src/`

(every algorithm:
VAD, diarize, LID, alignment, cache, registry) consumed by all
language wrappers, with thin presentation layers in `examples/cli/`

.
Per-model runtimes live in `src/{whisper,parakeet,canary,...}.cpp`

,
sharing primitives from `src/core/`

(mel, ffn, attention, GGUF
loader, FastConformer / Conformer / Granite-LLM blocks, etc.).

— full layered layout, file-by-file tour of`docs/architecture.md`

`src/`

and`examples/cli/`

, per-backend internals table, regression discipline.— adding a new backend in five files, clang-format-18 setup, the`docs/contributing.md`

`crispasr-diff`

PyTorch-ground-truth workflow, and the TTS audio-cosine-vs-reference regression target.—`docs/regression-matrix.md`

`tools/test-all-backends.py`

capability tiers, cache modes (`keep`

/`ephemeral`

),`--skip-missing`

for CI.

**Shared libraries** (cross-repo with CrispEmbed):

`crisp_audio/`

— Whisper-shape audio encoder (Conv-stem + Transformer)`crisp_punc/`

— punctuation restoration (FireRedPunc + PCS)`crisp_lid/`

— text-based language identification (fastText + CLD3)`crisp_truecase/`

— truecasing (statistical + CRF + BiLSTM)

Both are self-contained static libraries with CMakeLists.txt. CrispEmbed
links them via `add_subdirectory(../CrispASR/crisp_*/)`

; CrispASR uses
them directly. If the shared dir is absent, both repos fall back to local
copies of the source files.

For benchmarks see [ PERFORMANCE.md](/CrispStrobe/CrispASR/blob/main/PERFORMANCE.md); for the
session-by-session port log and the bug-class lessons, see

[.](/CrispStrobe/CrispASR/blob/main/LEARNINGS.md)

`LEARNINGS.md`

`build/bin/crispasr-quantize`

is a single, model-agnostic GGUF
re-quantization tool that works across all supported model families
(Whisper, Parakeet, Canary, Cohere, Voxtral, Qwen3, Granite, Wav2Vec2,
MiMo-ASR, GLM-ASR, Moonshine, VibeVoice, Kokoro, Qwen3-TTS, …):

```
./build/bin/crispasr-quantize input.gguf output.gguf q4_k
```

**See ** for the full guide:
supported quant types, K-quant alignment fallback, recommended quant
per backend, and worked examples for each architecture.

`docs/quantize.md`

All backends use `ggml_backend_init_best()`

which automatically picks the highest-priority compiled backend: CUDA > Metal > Vulkan > CPU. To force a specific backend:

```
# Force Vulkan even when CUDA is available
crispasr --gpu-backend vulkan -m model.gguf -f audio.wav

# Pin a specific GPU (useful on Vulkan systems with iGPU + dGPU)
crispasr --gpu-backend vulkan -dev 1 -m model.gguf -f audio.wav

# Force CPU (useful for benchmarking)
crispasr -ng -m model.gguf -f audio.wav

# CUDA unified memory (swap to RAM when VRAM exhausted)
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 crispasr -m model.gguf -f audio.wav
```

Build flags: `-DGGML_CUDA=ON`

, `-DGGML_METAL=ON`

, `-DGGML_VULKAN=ON`

.

Notes:

`--gpu-backend vulkan`

selects the Vulkan backend, but it does not choose which physical GPU to use. Use`-dev N`

to select the Vulkan device index.- On some Windows laptops, Vulkan device
`0`

is the Intel iGPU and the NVIDIA GPU is`1`

. If Vulkan looks unexpectedly slow, rerun with`-dev 1`

. - The Windows convenience script
`build-vulkan.bat`

creates a separate Vulkan-capable binary at`build-vulkan\bin\crispasr.exe`

.

For most backends, `-v`

/ `--verbose`

surfaces per-stage timings and
device picks. For headless / library use (where the CLI flag isn't
plumbed through), set `CRISPASR_VERBOSE=1`

instead.

```
# Per-stage timing breakdown (mel / encoder / prefill / decode):
crispasr -v --backend gemma4-e2b -m model.gguf -f audio.wav
# gemma4_e2b: mel 128x1099 (17.2 ms)
# gemma4_e2b: encoder done: 1536x275 (719.0 ms)
# gemma4_e2b: prefill done, first_token=3133 (1464.0 ms)
# gemma4_e2b: decoded 25 tokens (7748.3 ms total)
# crispasr: transcribed 11.0s audio in 7.75s (1.4x realtime)

# Hugging Face access for gated models (Voxtral, Gemma4-E2B, …):
HF_TOKEN=hf_xxx crispasr -m auto --backend gemma4-e2b -f audio.wav
```

The server has its own auth env: `CRISPASR_API_KEYS`

(see
[Server mode](/CrispStrobe/CrispASR/blob/main/docs/server.md)).

**Per-backend debug / bench / dump-dir env vars (developer)**

These are useful when porting a new backend or chasing a regression.
The `*_BENCH=1`

toggles emit per-stage timings even without `-v`

; the
`*_DEBUG=1`

toggles emit per-step diagnostic prints; the `*_DUMP_DIR=`

paths write per-stage F32 tensors for diff-testing against a PyTorch
reference (see [Debug a new backend against PyTorch ground truth](/CrispStrobe/CrispASR/blob/main/docs/contributing.md#debug-a-new-backend-against-pytorch-ground-truth)).

| Env var | Purpose |
|---|---|
`CRISPASR_VERBOSE=1` |
Forces verbose mode for any backend (parallel to the `-v` flag). |
`CRISPASR_DUMP_DIR=path/` |
Generic per-stage F32 tensor dump for the `crispasr-diff` harness. |
`GEMMA4_E2B_BENCH=1` |
Per-stage timings for the Gemma-4-E2B backend. |
`COHERE_BENCH=1` / `COHERE_DEBUG=1` |
Cohere transcribe per-stage timings / per-step diagnostics. |
`COHERE_PROF=1` |
Cohere graph-level profiling (per-op timings). |
`COHERE_THREADS=N` |
Override thread count for the Cohere backend. |
`COHERE_DEVICE=cpu|cuda|metal|vulkan` |
Force the Cohere backend onto a specific device. |
`COHERE_DUMP_ATTN=path/` |
Dump attention activations for Cohere (used by the diff harness). |
`FIRERED_BENCH=1` |
Per-stage timings for the FireRedASR backend. |
`FIREREDPUNC_DEBUG=1` |
Per-step diagnostics for the FireRed punctuation post-step. |
`MOONSHINE_STREAMING_BENCH=1` |
Per-stage timings for moonshine-streaming. |
`OMNIASR_BENCH=1` / `OMNIASR_DEBUG=1` / `OMNIASR_DUMP_DIR=` |
OmniASR per-stage timings, diagnostics, and stage dumps. |
`PARAKEET_DEBUG=1` |
Parakeet TDT per-step diagnostics (joint network, blank-id sanity). |
`QWEN3_TTS_BENCH=1` / `QWEN3_TTS_DEBUG=1` / `QWEN3_TTS_DUMP_DIR=` |
Qwen3-TTS per-stage timings, diagnostics, and stage dumps. |
`VIBEVOICE_BENCH=1` / `VIBEVOICE_DEBUG=1` / `VIBEVOICE_DUMP_DIR=` |
VibeVoice ASR per-stage timings, diagnostics, and stage dumps. |
`VIBEVOICE_REF_FEATURES=path` |
Replace the live encoder with a saved feature tensor (regression harness). |
`VIBEVOICE_TTS_DUMP=path/` |
VibeVoice TTS per-stage dumps (token IDs, base/TTS hidden, neg condition, frame-0 noise/v_cfg/latent/acoustic_embed) for the diff harness. |
`VIBEVOICE_TTS_DUMP_PERFRAME=1` |
Per-frame VibeVoice TTS dumps written as `perframe_<stage>_f<NNN>.bin` . Pair with `VIBEVOICE_TTS_DUMP=path/` and `VIBEVOICE_TTS_NOISE=path` for stage-by-stage AR diff against `tools/run_official_vibevoice.py` . |
`VIBEVOICE_TTS_TRACE=1` |
Extra one-line traces (negative-condition prefill rms, scaling/bias factors loaded). Same effect as library verbosity ≥ 2; there is no CLI flag for it (`-v` caps verbosity at 1). |
`VIBEVOICE_VOICE_AUDIO=path.wav` |
Reference voice WAV for 1.5B-base TTS without a `.gguf` voice cache. |
`VIBEVOICE_TTS_NOISE=path` |
Override the per-frame Gaussian init noise. Flat little-endian float32 `[N_frames, vae_dim]` — typically the `noise.bin` written by `tools/run_official_vibevoice.py` . |
`VIBEVOICE_VAE_BACKEND=cpu|metal|cuda|vulkan` |
Pin the VAE decoder onto a specific backend. |
`WAV2VEC2_BENCH=1` / `WAV2VEC2_VERBOSE=1` / `WAV2VEC2_DUMP_DIR=` |
wav2vec2 per-stage timings, verbose graph traces, and stage dumps. |
`CRISPASR_VOXTRAL4B_STREAM_TIMING=1` |
Per-stage timings for the voxtral4b streaming path (encoder drain / prefill / first-text-token / decode-step p50/p95). |
`CRISPASR_VOXTRAL4B_STREAM_CHUNK_MS=N` |
Override the internal encoder chunk size (default 240 ms). Must be a multiple of 80 ms. Larger = faster feed (kernel-launch amortisation), longer live-caption latency floor. |
`CRISPASR_VOXTRAL4B_STREAM_BATCH_ENCODER=1` |
Regression-debug: ignore the streaming encoder's audio_embeds and re-run the whole batch encoder at flush. |
`CRISPASR_VOXTRAL4B_STREAM_DEBUG=1` / `CRISPASR_VOXTRAL4B_STREAM_DIFF=1` |
Per-step decode prints / side-by-side encoder cosine vs the batch encoder. |
`CRISPASR_VOXTRAL4B_STREAM_LIVE=1` |
Live-captions decode-during-feed (PLAN #7 phase 3). `get_text()` polled during feed returns progressive transcript. Default OFF (PTT semantics). Wrappers: Python `Session.stream_open(live=True)` , Rust `stream_open_ex(.., live: true)` . |
`CRISPASR_VOXTRAL4B_STREAM_DECODER_THREAD=1` |
Decoder worker thread (PLAN #7 phase 4, implies live mode). Lets `feed()` return between encoder chunks without waiting for the decode loop — useful for mic-driven workloads. On M1 the Metal queue serializes encoder and decoder so total wall-clock is unchanged; faster GPUs with kernel-level parallelism see real overlap. |
`CRISPASR_VOXTRAL4B_FUSED_QKV=0` |
Opt out of the runtime fused-QKV LLM path (default-on, ~7-8 % decode speedup on M1 Q4_K, ~500 MB extra memory). |
`CRISPASR_QWEN3_ASR_FUSED_QKV=0` |
Opt out of the runtime fused-QKV LLM path for qwen3-asr (default-on; works on F16/F32/Q4_K/Q8_0/...). |
`CRISPASR_VOXTRAL_FUSED_QKV=1` |
Opt in to the runtime fused-QKV LLM path for voxtral 3B. Off by default (no measurable speedup on JFK-shape decodes; useful for long-form workloads where decode dominates). |
`QWEN3_TTS_FUSED_QKV=1` |
Opt in to the runtime fused-QKV talker path. |
`GRANITE_DISABLE_ENCODER_GRAPH=1` |
Force the granite-speech / -plus / -nar encoder back to the per-layer CPU loop (slower but kept around for debugging). The single ggml-graph encoder with per-layer Shaw RPE is the default and is bit-near-identical to the CPU loop while being ~2× faster end-to-end across all three variants. |
`CRISPASR_NO_REL_POS=1` |
Ablate the relative-position bias in the Gemma-4 audio encoder (development only). |
`ECAPA_REF_FBANK=path` |
Reference filterbank tensor for the ECAPA-TDNN LID model (regression harness). |
`CRISPASR_SHERPA_LID_BIN=path` |
Override the auto-detected sherpa-onnx LID binary. |
`CRISPASR_ARG_DEVICE=N` |
Default GPU device index when `-dev` isn't passed. |
`GGML_CUDA_ENABLE_UNIFIED_MEMORY=1` |
Let CUDA swap to RAM when VRAM is exhausted. |
`GGML_VK_VISIBLE_DEVICES` / `CUDA_VISIBLE_DEVICES` |
Standard ggml/CUDA device-visibility filters. |

`HF_TOKEN`

and `HUGGING_FACE_HUB_TOKEN`

are both honoured for gated-model
downloads (in that order).

— the original ggml inference engine and Whisper runtime this fork is built on[whisper.cpp](https://github.com/ggml-org/whisper.cpp)— the tensor library everything runs on[ggml](https://github.com/ggml-org/ggml)**NVIDIA NeMo**— parakeet-tdt-{0.6b-v2,0.6b-v3,1.1b}, parakeet-tdt_ctc-{110m,1.1b,0.6b-ja}, parakeet-ctc-{0.6b,1.1b}, canary-1b-v2, canary-ctc aligner, and the FastConformer-CTC family (stt_en_fastconformer_ctc_{large,xlarge,xxlarge} plus CTC branches of the stt_*_fastconformer_hybrid_large[_pc] fleet: en-pc, de, es, fr, it, nl, pl, ru, ua, hr, be, ar, fa, ka, hy, uz, kk-ru — all usable both as ASR backends and as compact ~82 MB`-am`

forced aligners)**Cohere**— cohere-transcribe-03-2026** Qwen team (Alibaba)**— Qwen3-ASR-0.6B, Qwen3-ASR-1.7B, Qwen3-ForcedAligner-0.6B** Mistral AI**— Voxtral Mini 3B and 4B Realtime** IBM Granite team**— Granite Speech 3.2-8b, 3.3-2b, 3.3-8b, 4.0-1b** Meta / wav2vec2**— wav2vec2 CTC models (XLSR-53 English, German, multilingual via any Wav2Vec2ForCTC checkpoint)— optional diarization via subprocess (ONNX models)[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)— VAD (native GGUF) and language identification (native GGUF, 95 languages)[Silero](https://github.com/snakers4/silero-vad)— speaker diarization segmentation (native GGUF port)[pyannote](https://github.com/pyannote/pyannote-audio)+[miniaudio](https://miniaud.io/)+[stb_vorbis](https://github.com/nothings/stb)— embedded/linked audio decoders (WAV/MP3/FLAC/AIFF/OGG/Opus, no ffmpeg; AAC/M4A/ALAC via Apple AudioToolbox)[libopus/opusfile](https://opus-codec.org/)(MIT) — in-tree clean-room codec suite: MP3/AAC-LC/Ogg-Opus output encoders (TTS[glint](https://github.com/CrispStrobe/glint)`.mp3`

/`.aac`

/`.opus`

) and cross-platform ADTS AAC-LC + Ogg Opus input decoders (no libopus needed)(Anthropic) — significant portions of the crispasr integration layer, all model converters, and the FastConformer/attention/mel/FFN/BPE core helpers were co-authored with Claude[Claude Code](https://claude.ai/claude-code)

Same as upstream whisper.cpp: **MIT**.

Per-model weights are covered by their respective HuggingFace model licenses (see [Supported backends](#supported-backends)). The `crispasr`

binary itself links model runtimes that are mostly permissively licensed (MIT / Apache-2.0 / CC-BY-4.0 for weights).
