Meta has introduced Muse Voice Transcribe, a new AI-powered speech-to-text model designed to transcribe conversations in real time. Developed by Meta Superintelligence Labs, the model can understand and process multiple languages within the same conversation, making it useful for multilingual speakers. Muse Voice Transcribe also marks Meta’s first real-time audio perception model from the newly formed AI research division.
Supports 5 Indian Languages #
Meta says Muse Voice Transcribe can work with more than 70 languages, including Hindi, Tamil, Telugu, Malayalam and Kannada. The AI model is also built to recognise code-switching, allowing users to move between languages during the same conversation without requiring separate models.
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
— Mark Zuckerberg (@finkd)[pic.twitter.com/LViMDSkbim][September 1, 2026]
Real-Time Multilingual Transcription #
The model generates text as people speak, rather than waiting for an entire recording to finish. It can also distinguish between speakers in recordings featuring more than 20 voices and process audio longer than an hour, with these functions handled within a single model.
Balancing Speed and Accuracy #
Muse Voice Transcribe is designed to adjust how long it listens before producing each word. It can respond quickly when speech is easy to understand while taking additional time to process words that are more difficult to recognise. Meta says the model topped the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026, while 25 of the more than 70 supported languages had been validated at launch.
Available Through Meta’s Model API #
Meta has made Muse Voice Transcribe available through its Model API, where it costs $3 per 1,000 audio minutes, or about $0.18 per hour according to the company. The technology is already being used for dictation in Meta AI for Mac and Muse Code, highlighting its potential for transcription, coding, voice assistants and other speech-based applications.
ALSO SEE: