Meta introduced Muse Voice Transcribe, a real‑time audio model capable of speaker diarization and multi‑language transcription for up to 20 speakers.

Meta has unveiled Muse Voice Transcribe, a real‑time audio model that can simultaneously transcribe and differentiate up to 20 speakers while supporting multiple languages.

Key capabilities of Muse Voice Transcribe

The model combines speaker diarization with automatic speech recognition, allowing it to assign spoken content to individual participants in a conversation. It works in real time, delivering transcription results as audio is captured.

Muse Voice Transcribe is designed for multilingual environments, automatically detecting and transcribing speech in several languages without requiring a separate language selection step.

Potential applications

Enterprises can use the technology for large‑scale meetings, webinars, and virtual events where many participants speak, reducing the need for manual note‑taking and improving accessibility for deaf or hard‑of‑hearing attendees.

Media organizations may employ the model to generate live captions for panel discussions or press briefings, ensuring accurate speaker attribution in real time.

Technical highlights

  • Speaker diarization for up to 20 concurrent voices
  • Support for multiple languages with automatic detection
  • Low latency processing suitable for live streaming
  • Integration options via Meta’s AI platform APIs

Meta says the model builds on its previous research in speech AI and leverages large‑scale training data to improve accuracy across diverse accents and acoustic conditions.

“Real‑time, multi‑speaker transcription opens new possibilities for collaboration and accessibility,” a Meta spokesperson said.

The company plans to roll out Muse Voice Transcribe to developers through its AI services suite, with documentation and sample code to facilitate integration into existing workflows.

For more details, see Dataconomy coverage of Meta’s Muse Voice Transcribe launch.