Meta has unveiled a new artificial intelligence model capable of transcribing long, real-time conversations involving multiple speakers and languages while identifying who said what.
Called MUSE Voice Transcribe, the model was introduced by Meta Superintelligence Labs on September 1. Meta described it as its first real-time audio perception model, combining speech-to-text transcription, speaker identification and detection of when one speaker stops and another begins.
The model can process conversations involving more than 20 speakers, making it suitable for meetings, interviews, discussions and other extended conversations.
MUSE Voice Transcribe uses streaming automatic speech recognition (ASR), converting speech into text as a conversation takes place rather than after the entire recording has been processed. Meta said the system can prioritise speed when speech is clear while spending more time analysing audio when words or context are ambiguous.
The model also supports speaker diarisation, allowing transcripts to be separated by speaker. Its endpointing capability detects when one speaker has finished and another has started, helping maintain the sequence of a conversation.
Meta said MUSE Voice Transcribe was trained on data covering more than 70 languages and tested on more than 25 languages. It can also detect code-switching, allowing it to handle conversations in which speakers switch between languages.
The model is currently available through Meta’s Model API and can also be used with Meta AI’s dictation feature for Mac and MUSE Code. Meta has priced access to the API at $3 per 1,000 minutes of audio.
The technology could support live transcription of online meetings and interviews, as well as applications in education, customer service and voice-based AI, particularly in settings involving multiple speakers or languages.