Meta’s Superintelligence Labs team has launched Muse Voice Transcribe, a real-time audio perception model that handles streaming speech recognition, speaker diarization, and sentence boundary detection. Processing audio in 80-millisecond chunks, the model dynamically balances latency and accuracy using reinforcement learning. In third-party evaluations by Artificial Analysis, the model achieved a 3.1% word error rate on English speech with a 0.16-second latency, placing it alongside competitive offerings from ElevenLabs and AssemblyAI.

The unified model supports over 70 languages, tracks up to 20 speakers simultaneously in a single session, and processes multi-language code-switching mid-sentence. Meta is integrating the capability into Meta AI and Muse Code, while offering public API access through the Meta Model API. The technology is positioned as foundational infrastructure for always-on ambient computing, including smart glasses.

Meta is pricing the service at $0.18 per hour ($3 per 1,000 audio minutes), heavily undercutting competing solutions from Cartesia, ElevenLabs, and Deepgram. While Meta is not releasing the model weights or training details, its aggressive pricing structure increases pressure on standalone speech-to-text API providers.

Why it matters

  • Undercuts industry speech-to-text API pricing by up to 50%+ compared to rival offerings from ElevenLabs, Cartesia, and Deepgram.

  • Combines low-latency transcription, multi-speaker diarization, and sentence boundary detection inside a single unified neural model.

  • Lays the technical software foundation for always-on ambient audio processing across future wearable hardware like smart glasses.

Source: the-decoder.com