Tech News
← Home  ·  All topics

Speaker Diarization

1 GoKawiil brief on this topic

Meta launches Muse Voice Transcribe, a real-time multilingual speech-to-text model

Meta unveiled Muse Voice Transcribe, its first real-time audio perception model, capable of transcribing over 20 speakers simultaneously while switching between more than 20 validated languages, including mid-sentence code-switching. CEO Mark Zuckerberg demoed the tool on X, noting it uses adaptive delay to balance speed and accuracy. The model is now available through Meta's Mac AI app, Muse Code, and its Model API at $3 per 1,000 audio minutes.