Meta has made a significant breakthrough in AI transcription technology with the introduction of Muse Voice Transcribe, its first real-time audio model. This innovative model can handle dictation and transcription for more than 20 speakers and can seamlessly handle multiple languages at once. Meta CEO Mark Zuckerberg recently shared an example of the model's ability to handle multiple speakers and languages simultaneously, showcasing its capability to automatically distinguish between speakers and switch between languages.
The model's advanced features include the ability to pick up on "code-switching" and transcribe sentences that use words from multiple languages. According to Zuckerberg, "The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy." This technology holds up well even on messy, real audio, and has been trained across over 70 languages, with 25 validated at launch.
Muse Voice Transcribe is the latest release from Meta's Superintelligence Lab (MSL), which has been actively developing new AI models and tools. The model is now available to developers within Muse Code and Meta's Model API, priced at $3 for 1,000 audio minutes. A demo version of Muse Transcribe is also available on Meta's research blog. Additionally, the model is already powering dictation in the Meta desktop app and Muse Code, and is live via Meta Model API.
This development comes less than a week after Google introduced its own audio model, Gemini 3.5 Transcribe, which boasts similar capabilities. While Google plans to integrate its audio model into Android and Chrome, it's unclear if Meta has plans to integrate Muse Voice Transcribe into its flagship services. For now, users can experience the new model's capabilities in Meta's recently released Meta AI Mac app, which can power voice-enabled features on other apps.
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.