Meta Releases Muse Voice Transcribe for Real-Time, Multi-Speaker Audio Meta Superintelligence Labs released Muse Voice Transcribe on September 1, a real-time audio model that transcribes speech, separates up to 20 speakers, and detects utterance endpoints, processing audio in 80-millisecond chunks. Priced at $0.18 per audio hour via Meta's Model API, the model was trained on 70+ languages with 25 validated, and Artificial Analysis ranked it first in a streaming speech-to-text comparison with a 3.1% word error rate, though the benchmark is primarily English-only. Meta Releases Muse Voice Transcribe for Real-Time, Multi-Speaker Audio - Muse Voice Transcribe processes audio in 80-millisecond chunks and supports streaming transcription, speaker labeling for more than 20 speakers and endpoint detection.