Industrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detection, streaming, and OpenAI-compatible API.
Self-hosted speech recognition with VAD, punctuation, and speaker pipelines.
You need one blanket licence — the toolkit is MIT but model licences vary.
About FunASR #
Industrial speech recognition. Up to 340x realtime, 26x faster than Whisper. 50+ languages. Speaker diarization · Emotion detection · Streaming · One API call
Quick Start · Colab · Benchmark · Model selection · Migration guide · Use cases · Deployment matrix · Models · Agent Integration · Docs · Contribute
No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.
FunASR is an open-source project written primarily in Python, with 20k stars on GitHub. It was last updated in August 2026.
pip install torch torchaudio