- Muse Voice Transcribe processes audio in 80-millisecond chunks and supports streaming transcription, speaker labeling for more than 20 speakers and endpoint detection. <sup>[1]</sup>
- Meta lists the model at $0.18 per audio hour through its Model API. <sup>[2]</sup>
- Artificial Analysis ranked it first in a streaming speech-to-text comparison, but the published result is primarily an English evaluation and does not establish performance across all 25 validated languages. <sup>[3]</sup>
Meta Superintelligence Labs released Muse Voice Transcribe on September 1, introducing a real-time audio model that transcribes speech, separates speakers and detects when an utterance ends. The model processes audio in 80-millisecond chunks and uses adaptive delay to trade a small amount of latency for better recognition on difficult words. [1]
The model is available through Meta Model API, Meta AI for Mac and Muse Code. Meta’s developer site lists pricing at $0.18 per audio hour, but the company has not said that Muse Voice Transcribe runs locally on a Mac or other device. [2]
A speech layer for Meta’s assistants #
Muse Voice Transcribe was trained on more than 70 languages, with 25 extensively validated for the initial release. Meta also says it supports code-switching within and between sentences, along with keyword and context biasing for names and other specialized terms. [1]
Artificial Analysis measured a 3.1% final-transcript word error rate and about 0.16 seconds from detected end of speech to the final result, placing the model first in its streaming comparison as of September 1. Independent coverage cautions that the benchmark is mainly an English streaming test, so the result does not verify Meta’s broader multilingual claims. [3]
Meta’s demonstrations show the model handling conversations with as many as eight speakers, while its documentation describes support for more than 20 speakers and audio sessions longer than an hour. The company has framed that capability as a foundation for agents that can follow real conversations through devices such as AI glasses. [1]
Consent remains part of the product design #
Meta says audio from its public transcription demo is processed to produce a transcript and is not stored. The demo also requires users to confirm that they have permission to record the captured audio. Those statements address the demo’s handling of recordings, but they do not by themselves define consent rules for third-party applications or future always-listening assistant features. [1]
Companies mentioned #
Further sources #
The stories that matter, in one email. Free — unsubscribe anytime.