Doubao Audio Generation Model 1.0 - An end-to-end audio creation model launched by Volcano Engine.
Doubao Audio Generation Model 1.0 is an end-to-end audio creation model launched by Volcano Engine, supporting text or audio as reference input to generate target audio. A single Prompt can be edited with multi-character dialogue, emotional tone, background music, and...
What is Doubao Audio Generation Model 1.0?
Doubao Audio Generation Model 1.0 is an end-to-end audio creation model launched by Volcano Engine, supporting text or audio as reference input to generate target audio. A single prompt can be arranged with dialogue from multiple characters, emotional tone, background music, and environmental atmosphere, directly producing complete audio works with narrative tension, eliminating the need for post-production multi-track mixing. The model maintains high timbre consistency during long-term generation, supports decoupled control of timbre and style, and covers scenarios such as audio dramas, podcasts, and brand audio.
Main functions of Doubao audio generation model 1.0
-
Reference generationIt supports text descriptions or reference audio as input and generates target audio end-to-end without additional training.
-
Full-element arrangementDefine character dialogue, emotional tone, background music, and environmental sound effects simultaneously in a single Prompt, and the output is the finished product.
-
Multi-role consistencyIt supports multi-character voice definition and long-term consistency maintenance, avoiding the "crossover" problem in long audio.
-
Nonverbal expressionIt accurately reproduces details such as laughter, sighs, pauses, and dialect accents, enhancing the vitality of the dialogue.
-
Tone style decouplingThe same timbre can be adapted to different emotions and scenarios, supporting differentiated expressions of "one voice, multiple perspectives".
-
Audio extensionBased on a 2-minute reference audio, the recording was extended multiple times to maintain a high degree of consistency in timbre.
Technical Principles of Doubao Audio Generation Model 1.0
- End-to-end multimodal generationThe model adopts a unified end-to-end architecture, encoding text descriptions and audio references into a shared latent space representation. The target audio waveform is directly generated through the decoder, avoiding the pipeline architecture of traditional TTS + sound effects + music track synthesis, and realizing the integrated generation of human voice, background music and ambient sound.
- Long-term tone consistency mechanismBy deeply linking the latent space features of the original audio and the reference audio, the timbre anchor point is locked during multiple audio extensions, ensuring that the character's voice features remain highly consistent between the 1st minute and the 10th minute, thus meeting the long-term generation needs of audiobooks, long dramas, etc.
- Decoupling control of timbre and styleThe model separates timbre identity features and emotional expression style into different subspaces, supporting flexible switching of the same speaker's timbre under different emotions and contexts, while realizing one voice with multiple roles, that is, the same voice base presents differentiated expressions under different role settings.
How to use Doubao audio generation model 1.0
Volcano ArkThe Doubao audio generation model 1.0 API is now open for invitation testing. Individual users can directly experience it at the Volcano Ark Experience Center https://ark.volcengine.com/region:cn-beijing/experience/voice?model=doubao-seed-audio-1-0&sessionid= and enjoy a 30-minute creation quota.
The core advantages of Doubao audio generation model 1.0
-
Integrated generation of all elementsSay goodbye to the tedious process of producing and editing vocals, sound effects, and music separately. A single Prompt can directly produce finished audio.
-
Long-term tone consistencyIt solves the core pain point of inconsistent character voices in long audio creation, and supports multiple extensions without the need for segment-by-segment audio editing.
-
Zero-sample multimodal creationIt supports dual-modal input of text and audio, and can generate high-quality target audio without additional training, greatly reducing the creative threshold.
-
Fine decoupling of timbre styleThe same timbre can be adapted to a variety of emotions and roles, enabling flexible "one voice, multiple roles" expression and enhancing the freedom of dubbing and performance.
Comparison of Doubao Audio Generation Model 1.0 with similar competing products
| Comparison Dimensions | Doubao Audio Generation Model 1.0 | AudioX-Turbo |
|---|---|---|
| Core positioning | End-to-end full-element audio creation (integrated vocals, music, and sound effects) | Multimodal audio generation and editing (text/image/video/audio → audio) |
| Input mode | Text description, reference audio | Text, image, video, and audio four modalities |
| Multi-role arrangement | A single Prompt supports unified arrangement of dialogue, tone, and emotion from multiple characters. | Primarily focused on single-audio generation, with limited ability to arrange long dialogues involving multiple characters. |
| Tonal consistency | Supports multiple extensions of long audio clips to maintain a high degree of consistency in the character's voice. | While it boasts strong single-generation capabilities, long-term consistency extensions are not explicitly supported. |
| Full element generation | Dialogue, background music, and ambient sound effects are output in one integrated output, eliminating the need for post-production mixing. | It can generate audio content, but its ability to integrate music, sound effects, and vocals into a single video is relatively weak. |
| Tone style decoupling | Supports matching the same timbre to different emotions and "one sound, multiple angles". | Style transfer is supported, but the character-level timbre decoupling control is relatively coarse. |
| Chinese optimization | Native Chinese context optimization, supporting dialect accents | It supports multiple languages, but its ability to express details in Chinese is slightly inferior. |
| Usage threshold | Prompt-driven, zero-sample creation, direct experience on Volcano Ark. | Requires a certain level of technical expertise; deployment is primarily via GitHub open-source platforms. |
Application scenarios of Doubao audio generation model 1.0
-
Audio dramas and podcastsCreators can directly generate complete audio works with multiple characters' dialogue, background music, and sound effects using Prompt, saving the need for post-production mixing.
-
Brand audio advertisingQuickly produce brand audio materials including narration, background music, and ambient sounds, shortening the advertising production cycle.
-
Long audio contentAudiobooks and long-running TV series utilize the timbre consistency extension function to maintain the character's voice throughout.
-
Live Streaming E-commerce AudioGenerates audio scripts with specific accents and emotional rhythms for product sales, adaptable to different products and broadcaster styles.
-
Film and television pre-dubbingIt can quickly generate temporary dialogue and ambient sounds for film and television clips to assist in pre-production editing and storyboard confirmation.