AB
AiBoss
News

Meta launches its first real-time audio perception model, Muse Voice Transcribe

Meta has launched its first real-time audio-aware model, Muse Voice Transcribe, which integrates streaming speech recognition, speaker segmentation, and endpoint detection. It supports transcribing over 20 speakers across multiple tracks and more than 70 languages. The model employs an adaptive latency mechanism, balancing real-time response and recognition accuracy, and ranks first on the Artificial Analysis streaming speech-to-text leaderboard.