Meta Superintelligence Labs has released Muse Voice Transcribe, a real-time speech-to-text model that streams transcription, detects sentence endpoints, and identifies more than 20 distinct speakers without a separate processing step. Trained on over 70 languages (25 thoroughly validated), the model handles long recordings beyond an hour and supports code-switching between languages, all offered via API at $0.18 per hour of audio.
venturebeat.com
· 2026-09-02
Meta unveiled Muse Voice Transcribe, its first real-time audio perception model, capable of transcribing over 20 speakers simultaneously while switching between more than 20 validated languages, including mid-sentence code-switching. CEO Mark Zuckerberg demoed the tool on X, noting it uses adaptive delay to balance speed and accuracy. The model is now available through Meta's Mac AI app, Muse Code, and its Model API at $3 per 1,000 audio minutes.
engadget.com
· 2026-09-01
Meta Superintelligence Labs has released Muse Voice Transcribe, a streaming speech recognition model that combines transcription, speaker diarization for over 20 voices, and endpointing in a single pass. It supports more than 70 languages, with 25 validated at launch, handles recordings over an hour long, and allows mid-sentence code-switching between languages. The model is live now in Meta AI for Mac, Muse Code, and via the Meta Model API at $3 per 1,000 audio-minutes.
9to5mac.com
· 2026-09-01