{"slug": "msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker", "title": "MSL rolls out Muse Voice Transcribe, a real-time audio model with speaker diarization", "summary": "Meta Superintelligence Labs has released Muse Voice Transcribe, a real-time speech-to-text model with speaker diarization and endpointing, available through Meta's Model API as part of the Muse Spark family. The model aims to improve transcription accuracy in multi-speaker settings by automatically attributing speech and detecting when a speaker has finished, though pricing and benchmarks have not been disclosed.", "body_md": "Photo: Tima Miroshnichenko / Pexels\n\n# MSL rolls out Muse Voice Transcribe, a real-time audio model with speaker diarization\n\nMeta Superintelligence Labs debuts its first product, a speech-to-text model that can tell who's talking and when they stop\n\nMeta Superintelligence Labs, the division Meta set up to chase artificial general intelligence, has shipped its first product. It’s not a reasoning engine or a world model. It’s a transcription tool.\n\nMuse Voice Transcribe is a real-time speech-to-text model that handles two tasks most transcription services still struggle with: speaker diarization (figuring out who said what) and endpointing (knowing when someone has actually finished talking versus just pausing to think). The model is now available through Meta’s Model API as part of the broader Muse Spark family.\n\n## What Muse Voice Transcribe actually does\n\nReal-time transcription sounds simple until you’ve tried to use it in a meeting with more than two people. Most existing tools either mash everyone’s words into a single undifferentiated stream or require manual speaker labeling after the fact. Diarization solves that by automatically attributing speech to individual speakers as it happens.\n\nEndpointing is the other half of the puzzle. It’s the system’s ability to detect when a speaker has genuinely finished a thought, as opposed to taking a breath or collecting themselves mid-sentence. Get this wrong and you end up with transcripts that chop sentences in half or lag behind the conversation by several seconds.\n\nThe Muse Spark family, which houses this model, is designed around natural conversation patterns. That means the models can handle interruptions, a feature that matters enormously for anything beyond scripted dictation. The system also supports multiple languages, though Meta hasn’t specified exactly which ones or how many.\n\nDevelopers can access Muse Voice Transcribe through Meta’s Model API, which entered public preview in mid-2026 with a pay-as-you-go pricing structure. The specific per-minute or per-token costs haven’t been disclosed yet.\n\n## The competitive landscape\n\nMeta isn’t entering an empty field. OpenAI’s Whisper set a high bar for open-source transcription quality. Google’s Chirp models power speech recognition across its cloud platform. Startups like AssemblyAI and Deepgram have built entire businesses around real-time transcription APIs with speaker identification.\n\nThe absence of independent benchmarks is worth noting. Without head-to-head comparisons on standard datasets like LibriSpeech or earnings call transcriptions, it’s hard to evaluate where Muse Voice Transcribe sits relative to existing options on raw accuracy. Specific pricing details, independent benchmarks, and comprehensive technical specifications for Muse Voice Transcribe have not been publicly released.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker", "canonical_source": "https://cryptobriefing.com/msl-muse-voice-transcribe-real-time-audio/", "published_at": "2026-09-01 17:19:47+00:00", "updated_at": "2026-09-01 17:25:13.746018+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-products", "ai-research"], "entities": ["Meta Superintelligence Labs", "Meta", "Muse Voice Transcribe", "Muse Spark", "OpenAI", "Whisper", "Google", "Chirp"], "alternates": {"html": "https://wpnews.pro/news/msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker", "markdown": "https://wpnews.pro/news/msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker.md", "text": "https://wpnews.pro/news/msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker.txt", "jsonld": "https://wpnews.pro/news/msl-rolls-out-muse-voice-transcribe-a-real-time-audio-model-with-speaker.jsonld"}}