Meta launches Muse Voice Transcribe, a real-time multilingual speech-to-text model
Meta unveiled Muse Voice Transcribe, its first real-time audio perception model, capable of transcribing over 20 speakers simultaneously while switching between more than 20 validated languages, including mid-sentence code-switching. CEO Mark Zuckerberg demoed the tool on X, noting it uses adaptive delay to balance speed and accuracy. The model is now available through Meta's Mac AI app, Muse Code, and its Model API at $3 per 1,000 audio minutes.