Meta Enters the Voice AI Race with Muse Voice Transcribe

Share
Meta Enters the Voice AI Race with Muse Voice Transcribe

Meta is stepping deeper into the rapidly evolving voice AI market with Muse Voice Transcribe, a real-time audio perception model developed by Meta Superintelligence Labs that combines speech recognition, speaker identification and endpoint detection in a single system.

The company says the model delivers streaming automatic speech recognition (ASR), diarization for more than 20 speakers and real-time endpointing. It is also designed to handle multilingual conversations and seamless code-switching, allowing speakers to move between languages without requiring separate models or additional processing.

The technology addresses a growing challenge as AI moves from chat interfaces to voice-based assistants and agents. Machines need to understand conversations in real-time, rather than simply transcribe audio after they are over.

Meta says Muse Voice Transcribe is trained on more than 70 languages, with 25 extensively validated for the initial release. The validated language set includes Hindi, while Meta’s India launch coverage says the model supports five major Indian languages — Hindi, Tamil, Telugu, Malayalam and Kannada.

One of its key features is adaptive delay. Rather than waiting the same amount of time before producing every word, the model dynamically determines how much audio context it needs, balancing transcription accuracy against latency. Meta says reinforcement learning is used to optimise that trade-off.

"The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy. With adaptive delay, the model is near the pareto front on speed-accuracy tradeoff," Mark Zuckerberg, Meta CEO said.

On Meta’s reported benchmarks, Muse Voice Transcribe ranked first on Artificial Analysis’ streaming speech-to-text leaderboard as of September 1, with a final-transcription word error rate of 3.1%. It also reported the lowest average diarization error rate across the AMI-IHM, AMI-SDM and VoxConverse benchmarks, at 17.5%.

Diarization is particularly important for applications involving multiple people. Muse can process audio exceeding an hour and distinguish more than 20 speakers without requiring post-processing, according to Meta.

The company has built diarization and end-pointing directly on top of its streaming ASR system. Special tokens allow the model to identify when a speaker changes and when speech begins or ends, enabling the system to understand the structure of a live conversation.

Meta is also positioning the model as an interface for its broader AI ecosystem. Voice dictation across Meta AI and Muse Code is now powered by Muse Voice Transcribe, and the model is available through Meta’s Model API, Meta AI for Mac and Muse Code.

Voice AI Space Heats Up

The move comes as voice AI in India is rapidly shifting from call automation to agentic systems. Sarvam has expanded from speech models into voice agents designed to conduct conversations, execute workflows and operate across Indian languages. The company says its broader AI stack now handles more than two million voice conversations every day.

Other startups are moving in the same direction. Ringg, which recently raised $10 million from Peak XV, processes about 20 million call attempts each month and is expanding beyond outbound calling into healthcare appointments, e-commerce recovery and fintech onboarding.

Gnani, meanwhile, recently launched Artha, a sovereign AI stack built around its 30-billion-parameter Evon 3.3 model and Plexus agentic platform, reinforcing the push toward Indian-language AI that can be deployed within domestic infrastructure.

The result is a market increasingly competing on more than voice quality. Latency, code-switching, Indian accents, speaker recognition, privacy, deployment and the ability to actually complete tasks are becoming the new battlegrounds.

For Meta, Muse Voice Transcribe is an important step toward its broader vision of personal AI that can listen to conversations in real-time. For India’s voice AI ecosystem, it is another sign that the technology is moving from simply answering calls to becoming an always-on interface between people, businesses and AI.