Doubao Audio Generation Model 1.0 Test - Multi-character voice acting, one-click audiobook generation
You might also be frustrated by this: AI-generated blockbusters have incredibly high-quality visuals, but the moment the characters start speaking, it instantly pulls you out of the story. The scene is clearly about a life-or-death situation, but the voiceover sounds like someone calmly reciting a product description...
You might also be troubled by: using AI The generated film has excellent visual quality, but the moment the character speaks, it instantly pulls the viewer out of the story.
The scene clearly depicts a life-or-death situation, yet the voice-over sounds like someone calmly reciting a product manual; any emotional fluctuations are achieved only through stiff shouting... What's worse, the character's voice is completely different in different segments, making it difficult to maintain consistency.
Not to mention that you also need to add ambient sounds, sound effects, background music, and lip-sync later on... It's a real hassle.
At the recent Volcano Engine FORCE conference, ByteDance officially released Doubao Audio Generation Model 1.0, which can now generate rich and emotional sound materials end-to-end.
We input a segmentPrompt wordsDoubao Audio Generation Model 1.0 can package and generate voice, sound effects, background music, and scene sounds all at once. It not only eliminates the tedious multi-track mixing and editing but also simulates the subtle breathing and emotional changes of a real person speaking. AI The voice sounds more natural and human..
How does it perform in real-world creative work? Let's put it to the test today.
We open the Volcano Ark Experience Center, select Doubao Audio Generation Model 1.0, and regular users have 30 minutes of audio...freeThe trial credit limit can be accessed via API later.
Our input effectPrompt wordsCombine the synthesized text and click "Generate" to get a complete audio clip containing human voices and ambient sounds.
Solo voice acting
I tried to generate a monologue from a character in a novel.
The background music lays a light foundation, featuring low strings, distant drumbeats, and ethereal female vocals. The opening is somber and oppressive, like the silence before a snowstorm. As the character's emotions progress, the music gradually intensifies, but never drowns out the vocals. The overall atmosphere shifts from a lone figure facing a predicament to breaking free and establishing a new path—solemn, tragic, restrained, yet intensely powerful. Xie Chang'an (a young female voice, clear and transparent, stable, slightly subdued, gradually becoming more resolute and powerful in the latter part) speaks calmly and restrainedly, as if establishing her own path under the watchful eyes of the masses: "All the nobles in the court cherish their own lives, so only a nobody like me can take action. My path is the path of all beings, a path that everyone can walk. Where there is a predicament, there is a way to break it. Rather than going with the flow, it's better to fight for survival. Perhaps a brighter future awaits on an unexpected third path."
At first, I wrote about young female voices, clear and ethereal, but the voice ended up being too soft; it had an ethereal quality, but lacked a sense of power. Later, I...Prompt wordsIf you change it to a young female contralto voice, without sweetness, softness, or girlishness, the effect will be very close to that of a leading lady.
The synchronously generated background music is also very powerful and matches the characters' voices and emotions very well.
Multi-character voice acting
We uploaded a crosstalk performance featuring two characters with very different personalities:
The background music is very soft, mainly featuring the opening gongs and drums of a small theater and short, crisp sanxian (a three-stringed plucked instrument). There are slight ambient sounds from the audience at the beginning, creating a lively, relaxed, and down-to-earth atmosphere. Laughter can occur sparingly, but not frequently, and should not drown out the dialogue. Vocals must be clear and prominent.
The female comedian (a young woman with a bright, clear voice, speaks quickly and fluently, with a touch of Beijing accent and a playful quality; her emotions are outgoing but not sharp) is excited and smug, saying as if she has discovered a new tool: "Let me tell you, now..." AI The voice-over is amazing; I just typed in the script, and it spoke it aloud for me.
The male straight man (a middle-aged man with a deep, resonant voice, speaking slowly and steadily, with a hint of dry humor and skepticism) calmly and questioningly replied, "What's so new about this? We've said this before."
The female comedian (a young woman's voice, raised in a high-pitched, exaggerated but cute tone) said, "Was that called talking before? It was called elevator announcements before."
The male straight man (middle-aged male voice, a beat slow, earnestly responding to the joke) said, "They're quite disciplined."
Female comedian (young female voice)fast(Catching it, with a smile) he said, "Discipline is there, but there's absolutely no affection."
The male straight man (middle-aged man, chuckles softly) says, "The main theme is equality for all beings."
The female comedian (a young woman's voice, still excited and speaking quickly) said, "It's different now. If you ask it to tell a children's story, it can be gentle; if you ask it to tell a suspenseful short drama, it can lower its voice; if you ask it to tell a story about a strong female lead, it can even bring a sense of breaking the deadlock."
The male straight man (middle-aged man, feigning doubt) said, "Then how about letting it perform crosstalk?"
The female comedian (a young woman's voice, pausing briefly, then speaking seriously) said, "Aren't we talking about this right now?"
The male straight man (middle-aged man, a beat slow to react, suddenly realizes) says, "So I was created too?"
In actual testing of two-person conversations, the naturalness of the conversations was much better than that of regular TTS.
The female lead has a faster pace and very natural emotional transitions, while the male supporting lead reacts more slowly. Each of their voices is very distinctive, and the consistency of their voices is also excellent.
The key point is that Doubao Audio Generation Model 1.0 can also directly generate the laughter of the audience at a crosstalk performance, which is very natural.
A single sentence can immerse you in the scene.AI The improvement in dubbing efficiency is evident.
Long audiobook text
Complex audiobooks often require the coordination of multiple characters and ambient sounds. We experimented with a complex ancient-style suspenseful ensemble piece:
The background music is subtly understated, featuring low strings, distant drumbeats, and a cool, mellow guqin, creating an overall atmosphere of solemnity, coldness, and oppression, with a sense of ancient political intrigue. In the first chapter, depicting the palace gates and court scene, the music is solemn and tense, like a snowstorm pressing down on a city. In the second chapter, depicting a secret meeting in a side hall, the music is even lower and darker, adding a touch of suspense. Ambient sounds include the sounds of wind and snow, the opening and closing of palace gates, the rustling of clothing, the crackling of lamp wicks inside the hall, and the distant footsteps of imperial guards. Voices must be clear and prominent, and the music and ambient sounds should not drown out the dialogue. The narrator (an adult female voice, deep and steady, with a strong narrative feel, a moderate to slow pace, and a vivid and suspenseful voice, avoiding a broadcasting tone) is calm and restrained, as if recounting a court intrigue unfolding on a snowy night. Shen Zhaoxue (a young female mezzo-soprano, with a cold, stable, and low voice supported by her chest cavity, clear enunciation, and clean endings; not sweet, soft, or girlish) is restrained, calm, and sharp, initially suppressing her anger, but gradually revealing her decisiveness and sense of control in breaking the deadlock. Xiao Cheng (a young male voice, deep and cold, speaking slowly and with restraint, carrying the aloofness and probing of the Crown Prince) is cautious, composed, and repressed, like someone who has been lying low for years testing a potentially dangerous knife. Pei Jingzhi (a middle-aged male voice, deep and cold, speaking slowly and with steady enunciation, carrying the oppressive and scrutinizing presence of a powerful minister) is composed, arrogant, and dangerous, like someone accustomed to controlling the court encountering an uncontrollable variable for the first time. The young emperor (a boyish male voice, somewhat immature but trying to be upright, with a tense and uneasy tone) is overwhelmed by the political situation, wanting to know the truth yet fearing it. Minister Zhou (middle-aged male voice, slightly weak, speech initially steady then disordered) is in a state of guilt, panic, and struggling to maintain composure. The Imperial Guard/Commander (adult male voice, low and short, tone obedient and tense) is solemn and vigilant. The young eunuch (young male voice, trembling, unsteady breathing) is terrified, on the verge of collapse, and desperate for survival. On the day Shen Zhaoxue arrived in the capital, the obituary from the northern border arrived before her. The obituary clearly stated: Shen Zhaoxue, the grain transport commissioner of the Zhenbei Army, encountered bandits while escorting military supplies; she and her cart plunged into the Black Gorge, leaving no trace. Yet, at dusk, she stood outside the Zhuque Gate, wearing a faded fox fur coat and leading a thin horse. The Imperial Guards guarding the gate saw the half-broken bronze sparrow talisman at her waist, and their faces changed instantly. The bronze sparrow talisman was a token bestowed by the late Emperor upon the Zhenbei Army for troop deployment; one half was in the northern border, and the other half was on the imperial desk. Everyone in the world knew that the other half of the bronze talisman from the northern border had disappeared ten years ago after the entire Shen family was imprisoned. Shen Zhaoxue raised her hand and placed the bronze talisman in the guard's palm. "Please announce my arrival," she said. "The dead have returned to the capital and wish to see the living officials." The wind and snow rushed into the palace gates, and the guard's hand trembled. Half an hour later, the Taiji Hall was brightly lit. The hall was filled with people. The Left Prime Minister, Pei Jingzhi, wore a purple robe, his ivory tablet tucked into his sleeve. He was over fifty years old, with thin eyelids, and when he looked up at people, it was as if he were looking at a piece of paper about to be burned. Crown Prince Xiao Cheng sat at the foot of the steps, his fingertips slowly caressing his teacup. The young emperor beside him was only twelve years old, his dragon robe on his shoulders looking as if it were borrowed. Shen Zhaoxue knelt in the hall, snow water dripping from the hem of her robe onto the blue bricks. Pei Jingzhi spoke first. "Since the guilty daughter of the Shen family is not dead, why doesn't she go to the Ministry of Justice to surrender herself?" Shen Zhaoxue raised her head. Her face was pale, but her eyes were steady. "If I go to the Ministry of Justice first, the esteemed officials will not hear any news from the Northern Border tonight." Someone in the hall sneered. "What news can a descendant of a disgraced official like you bring?" Shen Zhaoxue took out a roll of oilcloth from her sleeve and presented it with both hands. "170,000 shi of military rations left the Luocang three months ago, and the accounts say they have entered the Northern Border. But the Zhenbei Army only received 50,000 shi." The hall fell silent. Pei Jingzhi did not move. Crown Prince Xiao Cheng gently put down his teacup. "Continue," Shen Zhaoxue said, "The missing 120,000 shi, if converted into silver, would be enough to support 30,000 private soldiers for a year." Someone immediately rebuked, "Insolence! Do you know what you are saying?" "I know." Shen Zhaoxue looked at the man, "Minister Zhou, the Right Vice Minister of the Ministry of Revenue, the ink on the release document you approved was mixed with cinnabar. The half-grain token I picked up from Heixia also has this seal." The color drained from Vice Minister Zhou's face. Pei Jingzhi finally raised his eyes. “Miss Shen survived the fall, and she certainly has a silver tongue,” Shen Zhaoxue smiled. “Before the fall, I was also not one for words.” The wind outside the hall grew stronger. The young emperor gripped the armrest of his throne and asked softly, “What about the grain?” Upon hearing this, all the officials in the hall lowered their heads. Shen Zhaoxue looked at the young emperor. “The grain is gone.” She paused. “The northern border is almost gone too.” Crown Prince Xiao Cheng’s eyes darkened. “How is the Zhenbei Army?” “Seven days ago, the Qiang and Rong tribes breached the Frost River Pass. The Zhenbei Army retreated to Chensha City, where only two days’ worth of rations remain.” The young emperor stood up. “Why did no one report this?” Shen Zhaoxue did not answer immediately. She took out a second item from her bosom. A broken arrow. Half a piece of red cloth was wrapped around the shaft, the cloth now soaked black with blood. “Because the person who delivered the report died thirty miles before entering the capital.” She placed the broken arrow on the ground. “This is the sixth.” No one in the hall laughed anymore. Crown Prince Xiao Cheng slowly rose and descended the imperial steps. He stopped three steps away from Shen Zhaoxue, his gaze falling on the patch of unmelted snow on her shoulder. "What do you want?" "Open the granary." "Just open the granary?" "And a troop of imperial guards to escort me to Luo Cang to collect grain." Pei Jingzhi finally chuckled. "You want soldiers?" Shen Zhaoxue looked at him. "Prime Minister Pei is mistaken, I want the road." Pei Jingzhi's smile faded. "Luo Cang is in the capital region, and the guards are all under the command of the Ministry of Revenue. What right does the daughter of a disgraced official have to open the granary?" Shen Zhaoxue reached into her sleeve. The imperial guards all drew their swords. What she pulled out was a letter written in blood. Most of the words on the letter were blurred, only the last line was still legible. "Your subject, Shen Huaishan, is willing to exchange the lives of his entire family for three years of peace in the northern border." Shen Huaishan was her father. Ten years ago, he was accused of colluding with the Qiang and Rong tribes, and his entire family was imprisoned. Shen Zhaoxue was fifteen years old that year, kneeling at the entrance of the Ministry of Justice for three days, no one daring to give her a drop of water. Now, the blood-written letter, never delivered to the Emperor, lay on the throne, like a belated bone. The young Emperor's face paled. Pei Jingzhi's fingers twitched in his sleeve. Shen Zhaoxue saw it. She bowed her head, her voice low, yet it drowned out the wind and snow outside. "This humble woman relies on a memorial that the Shen family failed to deliver ten years ago, on the lives of seventy thousand soldiers in the Northern Frontier, and on the still-breathing civilians in Chensha City." She raised her head. "If that is still not enough, this humble woman is willing to sign a military pledge." Xiao Cheng asked, "How many days?" "Three days." "If the grain does not reach Chensha City?" Shen Zhaoxue looked at him and said, word by word, "I will die at the city gate." The hall was so quiet that one could hear the crackling of the lamp wick.
Doubao Audio Generation Model 1.0 will...automaticIt can identify audiobook content, such as descriptions of snowstorms entering the palace gates.automaticTo deduce and match suitable sound effects.
The female lead's voice is calm and restrained, while the minister's voice is slow and imposing. The narrator's voice and the voices of different characters are all highly recognizable.
The volume ratio of human voices, ambient sounds, and background music is also relatively moderate, saving us the tedious step of repeatedly adjusting the volume bar in editing software.
However, Doubao's audio generation model 1.0 can only generate a maximum of 2 minutes of audio at a time. To create a complete audiobook, it needs to be generated in segments.
The long text generation performance is generally poor, the order of some dialogues is reversed, and the recognition of polyphonic characters is not very stable, so the pronunciation needs to be noted.
AI Short drama dubbing
Let's try something more relatable. AI Short dramas. Regular TTS can only read the lines, but short dramas require the voice to have a sense of space.
The background music provides a subtle foundation, primarily featuring warm piano, gentle strings, and faint urban ambient sounds. The overall atmosphere is realistic, lifelike, with a touch of warmth and unexpected twists, avoiding suspense and horror. Ambient sounds include faint voices in the coffee shop, the clinking of cups, the doorbell, the vibration of a cell phone, and the sounds of vehicles on the street after the rain. Voices must be clear and prominent, and the music should not drown out the dialogue. The narrator (an adult female voice, gentle and steady, at a moderate pace, with a narrative feel) is calm and delicate, as if telling a small story that happened to an ordinary person. Lin Xia (a young female voice, clean and bright, with a slightly tired but restrained tone) transitions from disappointment and forced composure to gradual relief in the latter half of the story. Zhou Yan (a young male voice, deep and gentle, at a slow pace, sincere but slightly awkward) is cautious, guilty, and trying to explain, avoiding a domineering CEO tone. The shop assistant (a young female voice, light and natural, polite) appears briefly and naturally. Chapter Excerpt: "The Window Seat" Narrator: "Lin Xia and Zhou Yan met at that coffee shop, seven days after their breakup." Narrator: "The rain had just stopped, and the leaves outside the window were still dripping. Lin Xia sat by the window, two cups of coffee on the table. One was hot, the other was cold." Waitress: "Hello, would you like me to get you a hot one?" Lin Xia: "No, thank you." Narrator: "After she finished speaking, she glanced at her phone. Zhou Yan was twenty-six minutes late." Narrator: "When the wind chimes rang at the door, Lin Xia had already said, 'Don't contact me again.'" "I rehearsed it three times in my mind." Zhou Yan: "I'm sorry I'm late." Lin Xia: "You're always late." Zhou Yan: "There was really a traffic jam today." Lin Xia: "Last time it was overtime, and the time before that was a last-minute meeting. Zhou Yan, I'm not here to hear excuses." Narrator: "Zhou Yan stood by the table, holding a paper bag. The bag opening was slightly damp from the rain." Zhou Yan: "I know." Lin Xia: "Then sit down and finish what you have to say." Narrator: "He sat down opposite her, but didn't touch the cup of coffee that had gone cold." Zhou Yan: "You said that day..." I never put you first. Lin Xia: "Isn't it?" Zhou Yan: "Yes." "Narration:" Lin Xia raised her eyes and looked at him. This answer was too straightforward, and instead made the reproach she had prepared get stuck in her throat. Zhou Yan: "I always feel that if we do a good job first, save enough for the mortgage first, and stabilize our lives first, we will be better off." Lin Xia: "But what I've been waiting for is your absence again and again." Zhou Yan: "So I'm not here to ask for your forgiveness today." Lin Xia: "Then what are you doing here?" "Narration: "Zhou Yan pushed the paper bag in front of her. Zhou Yan: "Give you back your things." "Narration:" Lin Xia opened the paper bag. It wasn't the scarf she left at his house, nor was it the key. Narrator: "It's a stack of bus tickets, movie ticket stubs, and a dozen takeaway receipts." Lin Xia: "What is this?" Zhou Yan: "You said I don't remember anything." Actually I remember, I just didn’t say it. "Narration: "Lin Xia turned to the bottom and saw a faded sticky note. "Narration: "The above is what she wrote two years ago: If you have a quarrel in the future, go to the window and make up. "Lin Xia didn't speak. Zhou Yan: "I know it's a bit late to say this now." Lin Xia: "It is late." Zhou Yan: "Yeah." Narrator: "A car passes by outside the window, splashing water gently." Zhou Yan: "But I want to give them back to you. Not to make you turn back, but to tell you that I haven't forgotten those days." Lin Xia: "Then why didn't you say it sooner?" Zhou Yan: "Because I always thought that actions speak louder than words." Lin Xia: "And then?" Zhou Yan: "Then I realized that just doing without saying anything can also make people feel unimportant." Narrator: "Lin Xia looks down at the sticky note. The corner of the paper is curled up, but the words are still clear." Lin Xia: "Zhou Yan, I don't want to wait for someone who is always late anymore." Zhou Yan: "I know." Lin Xia: "But I can finish this cup of coffee with you." Narrator: "Zhou Yan paused for a moment, then slowly smiled." Zhou Yan: "It's cold." Lin Xia: "Then let's get a hot one." Narrator: "The clerk comes over and takes away the cold coffee. The clouds outside the window disperse a little, and sunlight falls on the window seat." Ending sound effects: The cup is gently set down, the doorbell rings once, and the background music gently fades away.
The characters' dialogue is very natural, allowing the viewer to feel the flow of emotions. The sounds of rain and cards turning help us create the scene.
Sound is no longer an accessory that is added at the end after the video is finished, but can be involved in the creation from the script stage.
Replica sound
Doubao Audio Generation Model 1.0 currently generates a maximum of 2 minutes of audio per session. If we want to create longer audio clips or sequels, how can we ensure the audio doesn't clash with the storyline?
We can upload reference audio or use previously generated audio as reference audio, with a maximum of 3 references supported at a time.Prompt wordsIt specifies that a certain character should use a certain timbre.
For example, let's try to replicate Doubao's voice:
The music opens with upbeat jazz drums, short bass notes, and a few playful piano riffs, accompanied by background sounds of hushed conversations, clinking glasses, and sporadic laughter from the audience in a small theater. The overall atmosphere is relaxed, lively, and reminiscent of an urban nightclub stand-up comedy show. Once the performer begins, the music quickly drops, leaving only a very light bass rhythm. Audience laughter, cheers, and applause are allowed to occur naturally, but should not drown out the vocals.
A stand-up comedian (young female voice, speaking Mandarin, low-pitched, slightly hoarse, medium to fast speaking speed, strong rhythm of witty remarks, with natural and punchline pauses, not in a broadcasting tone, performed by [actress's name]) speaks in a relaxed, self-deprecating manner, as if chatting with the audience in a small theater: "I recently discovered that..."AI The biggest impact wasn't finding a replacement job, but that it finally confirmed for my mother that I really am useless.
The audience chuckled.
The stand-up comedian (carefully setting the stage) continued, "My mom used to call me whenever she had a problem. She'd call me if her phone broke, if the TV had no sound, or if she couldn't find a WeChat group. Now it's different. She asks first..." AI.
Pause for half a second.
The stand-up comedian (his tone suddenly lowered) said, "After asking..." AI"Call me again."
The audience laughed.
The stand-up comedian said (helplessly): "She said,AI She gave me the answer, but she wasn't reassured and wanted me to confirm. I said, "Mom, you're putting me in a relaxed, lively, urban nightclub stand-up comedy vibe. The music starts with upbeat jazz drums, short bass lines, and a few playful piano notes, with background noise like hushed conversations, clinking glasses, and occasional laughter from the audience in a small theater. The overall atmosphere is relaxed, lively, and has an urban nightclub stand-up comedy feel. After the comedian starts speaking, the music drops rapidly, leaving only a very light bass rhythm. Audience laughter, cheers, and applause can occur naturally, but shouldn't drown out the voices. The stand-up comedian (a young female voice, speaking Mandarin, with a low pitch, slightly hoarse voice, medium to fast speaking speed, strong witty banter, with natural pauses and punchline pauses, not in a broadcasting tone, the performer is @audio1) is relaxed, self-deprecating, and says as if chatting with the audience in a small theater: 'I recently discovered...'"AI The biggest impact wasn't the replacement job, it was that it finally confirmed for my mom that I really am useless.” The audience chuckled. The comedian (earnestly setting the stage) continued, “Before, my mom would call me when she had a problem. Her phone would break, the TV wouldn't have sound, she'd miss a WeChat group. Now it's different, she asks first…” AI"..." After a half-second pause, the stand-up comedian (his voice suddenly lowered) said, "After asking..." AI"Call me again." The audience laughed. The comedian (helplessly) said, "She said,AI She gave me the answer, but she wasn't satisfied and wanted me to confirm it. I said, "Mom, you've demoted me from technical support to manual review." The audience's laughter intensified. The comedian (speech faster) said, "The scariest thing is, she now uses..." AI She used to post on her WeChat Moments: "Making dumplings today." Now it's: "Time settles in the flour, family affection shines in the wrinkles." A pause. The comedian (lowering his voice) said, "My dad saw it and asked her, 'Are these dumplings edible, or are they for exhibition?'" The audience laughed. The comedian (continued to tease) said, "My mom even asked me very seriously if posting like this would be too ordinary. I said no, it's fine, just not like you. She said, 'How not?' I said, 'You usually don't even use punctuation in your Moments, and suddenly family affection shines in the wrinkles, your relatives will think you've been possessed by flour.'" The audience laughed.
The generated timbre has a high degree of similarity to the reference timbre, and retains the self-deprecation and relaxed feel required for stand-up comedy. The pauses at punchlines and the interspersed laughter from the audience are very natural.
Doubao Audio Generation Model 1.0 can not only clone timbres, but also incorporate more emotions, making it more like performing a show using timbres.
In the past AI In voice acting, we simply feed the text to it; now, we need to...Prompt wordsLike a director, he instructs his characters—clearly describing their age, vocal characteristics, current emotions, movements, and background noise. The more concrete the details provided, the closer the final result will be to the desired outcome.
The previously tedious workflow of dubbing, music composition, finding sound effects, and mixing can now be streamlined through a reasonable process. Prompt fastOnce the first complete prototype is produced, the efficiency improvement is obvious. The production speed of short dramas, advertisements, courses, and virtual IPs will increase significantly.
Currently, the Volcano Ark Experience Center has opened up the experience of Doubao Audio Generation Model 1.0, and ordinary users can get 30 minutes of audio.freeExperience the benefits. In the future, it will also integrate with everyday tools such as CapCut and Tomato Novel, further lowering the barrier for ordinary people to create audio content.
If we talk about the past AI Dubbing solves the problem of whether there is sound, while Doubao Voice Model 1.0 started to solve the problem of whether the sound is engaging.
Of course, as version 1.0, Doubao Audio Generation Model 1.0 still has room for refinement and optimization in some complex physical sound field changes, polyphony, and accent details. However, the end-to-end generation potential demonstrated by Doubao Speech Model 1.0 has already shown us the nascent form of a revolution in audio productivity.
When images, videos, text, and audio AI The toolchain is becoming more and more complete.AI Voice-over will also be a key element in enhancing the content experience.
Original link:The sound went from "audible" to "expressive".AI The voice acting has really improved this time.