Universal-1 - A multilingual speech recognition and conversion model launched by AssemblyAI.
Universal-1 is a multilingual speech recognition and transcription model launched by AI speech startup AssemblyAI. It has been trained on more than 12.5 million hours of multilingual audio data and supports languages such as English, Spanish, French, and German.
What is Universal-1?
Universal-1 is a multilingual speech recognition and transcription model developed by AI speech startup AssemblyAI. Trained on over 12.5 million hours of multilingual audio data, it supports languages including English, Spanish, French, and German. This model delivers high-accuracy speech-to-text services in various environments, including noisy backgrounds, different accents, and natural conversations, and features fast response times and improved timestamp accuracy. Universal-1 is designed to improve the accuracy of every aspect of speech recognition, meeting customers' needs for nuanced speech data and serving as a powerful tool for building next-generation AI products and services.
Key features of Universal-1
- Multilingual supportUniversal-1 can handle multiple languages, including English, Spanish, French, and German, and has been optimized for these languages to improve the accuracy of speech recognition.
- High accuracyUniversal-1 maintains excellent speech-to-text accuracy under various conditions, such as background noise, accent diversity, natural dialogue, and language variations.
- Reduce hallucination rateCompared to Whisper Large-v3, Universal-1 reduces the illusion rate of speech data by 30%, meaning it reduces the number of times the model incorrectly generates text when there is no audio input.
- Rapid ResponseUniversal-1 is designed with efficient parallel inference capabilities, enabling it to quickly process long audio files and provide fast response times. Its batch processing capabilities are 5 times faster than Whisper Large-v3.
- Accurate timestamp estimationThe model provides timestamps accurate to the word level, which is crucial for applications such as audio and video editing and meeting recording. Universal-1's timestamp accuracy is 26% higher than Whisper Large-v3.
- User preferencesIn user preference testing, users preferred Universal-1's output 71% of the time, indicating that it better meets users' needs in actual use.
Universal-1 performance comparison
- Accuracy of English speech-to-text:Universal-1 achieved the lowest word error rate (WER) in 5 out of 11 datasets, compared to models such as OpenAI's Whisper Large-v3, NVIDIA's Canary-1B, Microsoft Azure Batch v3.1, Deepgram Nova-2, and Amazon and Google's Latest-long.
- Non-English speech-to-text accuracy:In tests on Spanish, French, and German, Universal-1 achieved lower WER scores in 5 out of 15 datasets, demonstrating its competitiveness in these languages.
- Timestamp accuracy:In terms of timestamp accuracy, Universal-1 improved the proportion of predicted words with timestamps within 100 milliseconds by 25.5% compared to Whisper Large-v3, increasing it from 67.2% to 84.3%.
- Reasoning efficiency:On an NVIDIA Tesla T4 machine, Universal-1 is 3 times faster than the faster whisper backend without parallelization, and it takes only 21 seconds to transcribe 1 hour of audio in 64 parallelized inferences.
- Reduced hallucinations:Universal-1 reduced the hallucination rate by 30% when transcribing audio compared to Whisper Large-v3.
- Human Preference Test:In human preference tests, evaluators preferred the output of Universal-1 in 60% of cases, compared to only 24% for Conformer-2.
- Voiceprint segmentation and clustering:Universal-1 achieves the following improvements in speaker diarization accuracy compared to Conformer-2:
- The Diarization Error Rate (DER) decreased by 7.7%.
- The combined measurement of WER and speaker marking accuracy reduced cpWER by 13.6%.
- The accuracy of speaker number estimation improved by 71.3%.
How to use Universal-1
Currently, Universal-1 is available in English and Spanish, with German and French versions coming soon. AssemblyAI will also add additional language support to future universal models. Interested users can try it out on the Playground or via the API.
- Try it out through Playground:The simplest way to try Universal-1 is throughAssemblyAI's Playground.In Playground, users can directly upload audio files or enter YouTube links, and the model will quickly generate a text transcription.
- Free API Trial:userCanFree registrationAnd obtain an API token.After registering, go to AssemblyAI's documentation (Docs) or Welcome Colab, which can help you get started with the API quickly.
For more information about Universal-1, please refer to AssemblyAI's official technical report:https://www.assemblyai.com/discover/research/universal-1
Application scenarios of Universal-1
- Dialogue Intelligence PlatformIt can quickly and accurately analyze large amounts of customer data, providing key customer voice insights and analysis, regardless of accent, recording conditions, or number of speakers.
- AI NotepadGenerates highly accurate, illusion-free meeting minutes, providing a foundation for the generation of summaries, action items, and other metadata based on large language models, including accurate proper nouns, speaker and time information.
- Creator Tools: Build AI-driven video editing workflows for end users, leveraging accurate speech-to-text output in multiple languages, low error rates, and reliable word timing information.
- Telemedicine PlatformAutomated clinical record entry and claim submission processes, utilizing accurate and faithful speech-to-text output, including rare words such as prescription names and medical diagnoses, with high success rates even under adversarial and far-field recording conditions.