Meta Superintelligence Labs releases Muse Voice Transcribe
Meta Superintelligence Labs releases Muse Voice Transcribe ASR model
Muse Voice Transcribe combines streaming ASR, speaker diarization, and endpointing in a single model rather than requiring separate pipeline components for each task. Meta is distributing it across multiple products and APIs.
The full picture
Meta Superintelligence Labs released Muse Voice Transcribe, a real-time audio perception model for streaming speech-to-text. The model handles speaker diarization with more than 20 speakers, turn-endpointing, and multilingual speech with code-switching, all within a single model. It processes audio in 80ms chunks and dynamically decides after each chunk whether to emit text or continue listening. Meta reports a 3.1% final-transcription word error rate and first-place rankings on Artificial Analysis streaming speech-to-text and public diarization benchmarks. Accuracy can be improved through language, keyword, and context biasing. The model is available through the Meta Model API, Meta AI for Mac, and Muse Code.
How it developed
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free