OpenAudio S1 - Fish Audio's next-generation speech generation model
OpenAudio S1 is a text-to-speech (TTS) model from Fish Audio, trained on over 2 million hours of audio data and supporting 13 languages. It employs a dual-autoregressive (Dual-AR) architecture and combines reinforcement learning with human feedback...
What is OpenAudio S1?
OpenAudio S1 is a text-to-speech (TTS) model from Fish Audio, trained on over 2 million hours of audio data and supporting 13 languages. Employing a dual autoregressive (Dual-AR) architecture and reinforcement learning with human feedback (RLHF) technology, it generates highly natural and fluent voices, almost indistinguishable from human voice acting. The model supports over 50 emotion and intonation markers, allowing users to flexibly adjust speech expression via natural language commands. OpenAudio S1 supports zero-shot and few-shot speech cloning, generating high-fidelity cloned voices with only 10 to 30 seconds of audio samples.
Main functions of OpenAudio S1
-
Highly natural voice outputBased on training with over 2 million hours of audio data, the generated voice is almost indistinguishable from human voice acting, making it suitable for professional scenarios such as video dubbing, podcasts, and game character voice acting.
-
Rich emotional and tone controlIt supports more than 50 emotion markers (such as anger, happiness, sadness, etc.) and tone markers (such as rapid, low, scream, etc.), and users can control the emotion and tone of voice through simple text commands.
-
Powerful multilingual supportIt supports up to 13 languages, including English, Chinese, Japanese, French, and German, demonstrating its powerful multilingual capabilities.
-
High-efficiency voice cloningSupports zero-sample and few-sample speech cloning, generating high-fidelity cloned voices with just 10 to 30 seconds of audio samples.
-
Flexible deployment optionsTwo versions are available: the full S1 with 4 billion parameters and the S1-mini with 500 million parameters. The latter is an open-source model suitable for research and educational use.
-
Real-time application supportUltra-low latency (less than 100 milliseconds) is ideal for real-time applications such as online games and live streaming content.
The technical principles of OpenAudio S1
-
Dual-AR architectureThis approach combines fast and slow Transformer modules to optimize the stability and efficiency of speech generation. The fast module quickly generates initial speech features, while the slow module fine-tunes these features to ensure naturalness and fluency of the speech.
-
Grouped Finite Scalar Vector Quantization (GFSQ) techniqueImprove code processing capabilities to reduce computational costs and increase model efficiency while ensuring high-fidelity voice output.
-
Reinforcement Learning and Human Feedback (RLHF)Through online RLHF technology, the model can more accurately capture the timbre and intonation of speech, resulting in more natural emotional expressions. Users can achieve subtle emotional control by tagging emotions such as (excitement), (nervousness), or (joy).
-
Large-scale data trainingTrained on an audio dataset of over 2 million hours, covering a wide range of language and emotional expressions, the model is able to generate highly natural and diverse speech.
- Voice cloning technologySupports zero-sample and few-sample speech cloning, requiring only 10 to 30 seconds of audio samples to generate high-fidelity cloned voices.
OpenAudio S1 project address
- Project official website:https://openaudio.com/blogs/s1
Application scenarios of OpenAudio S1
-
Content creationProvides professional-grade voiceovers for videos, podcasts, and audiobooks, significantly improving production efficiency.
-
Virtual AssistantCreate personalized voice navigation or customer service systems that support multilingual interaction and enhance user experience.
-
Games and EntertainmentGenerate realistic dialogue and narration for game characters to enhance player immersion.
-
Education and TrainingUsed to generate multilingual learning content to help students better understand and learn the pronunciation and intonation of different languages.
-
Customer service and supportSuitable for customer service robots, providing fast and accurate voice responses to improve the efficiency and quality of customer service.