Dubbing v2 - An AI dubbing model developed by ElevenLabs
Dubbing v2 is an AI dubbing model developed by ElevenLabs that supports automatic translation and dubbing of 29 languages, preserving the original speaker's timbre and emotion. The model offers a dual workflow mode: Auto Dub for quick preview generation, and Dub...
What is Dubbing v2?
Dubbing v2 is an AI dubbing model developed by ElevenLabs that supports automatic translation and dubbing of 29 languages, preserving the original speaker's timbre and emotion. The model offers a dual workflow: Auto Dub for quick preview generation, and Dubbing Project for segment-by-segment refinement in a timeline editor. Dubbing v2 supports multi-speaker separation, voice cloning, multi-format import/export, and API batch processing, capable of handling up to 2.5 hours of content.
Main features of Dubbing v2
-
AI-generated voiceoverSupports 29 languages, automatically detects multiple speakers and separates their voices, preserving the original sound characteristics.
-
Voice cloningIt offers three modes: fragment-level cloning, track-level cloning, and voice library selection.
-
Timeline EditorIt allows for segment-by-segment editing of transcribed text, adjustment of translation, fine-tuning of the timeline, and regeneration of fragments.
-
Multi-format supportImport supports MP3/MP4/WAV/MOV and YouTube/TikTok/Vimeo/X links; export supports MP4/AAC/WAV/SRT/AAF.
-
Dual workflow modeAuto Dub generates quickly and automatically; Dubbing Project supports detailed editing.
-
API IntegrationSupports batch processing and automated workflows, and can process up to 2.5 hours of content.
The technical principles of Dubbing v2
- Multilingual speech recognitionA deep learning-based ASR model automatically transcribes source language content, identifies multiple speakers, and separates audio tracks.
- Neural Machine TranslationIt employs a context-aware translation engine to preserve colloquial expressions and cultural context, avoiding distortion through literal translation.
- Speech cloning and synthesisThe speaker's vocal timbre features are extracted using a Speaker Encoder and combined with a TTS model to generate target language speech while preserving the original rhythm and emotion.
- Timeline alignment algorithmThe dynamic programming algorithm matches the translated text with the original timestamp, supporting segment-by-segment fine-tuning and regeneration.
- Multimodal processing pipelineAudio and video separation → speech recognition → translation → speech synthesis → mixed output, supporting continuous processing for up to 2.5 hours.
How to use Dubbing v2
- Visit the official websiteVisit the Dubbing v2 official website https://elevenlabs.io/dubbing-studio and log in to your ElevenLabs account.
- Upload source fileUpload MP3/MP4/WAV/MOV files directly, or paste the YouTube/TikTok/Vimeo/X platform link.
- Select target languageMultiple target languages can be selected for parallel processing at the same time.
- Select WorkflowAuto Dub quickly generates a preview, or Dubbing Project enters a fine editing mode.
- Review and EditingIn the timeline editor, check the translation accuracy segment by segment, adjust the timeline alignment, and regenerate unsatisfactory segments.
- Export finished productChoose to download in MP4 (with video), AAC/WAV (pure audio), or SRT subtitle format.
Dubbing v2's core advantages
-
High sound fidelityThe cloned voice-over is highly consistent with the original speaker's timbre, and the emotional expression is natural.
-
More people supportAutomatically identifies and separates different speakers, even when dialogue overlaps.
-
High edit controllabilityThe timeline editor offers segment-by-segment refinement capabilities, rather than an "all or nothing" output.
-
Cost efficiencyTraditional voice-over for a single 30-second ad in 10 languages can cost between $10,000 and $30,000, while ElevenLabs can complete it in minutes at a significantly lower cost.
Dubbing v2 project address
- Project official websitehttps://elevenlabs.io/dubbing-studio
Comparison of Dubbing v2 with similar competing products
| Dimension | Dubbing v2 | Speech Synthesis |
|---|---|---|
| Core Functions | Video/audio translation + dubbing + voice cloning | Text-to-speech with multiple voice options |
| Translation skills | Built-in automatic translation of 29 languages | No translation function |
| Tone Preservation | Preserve the original speaker's timbre and emotion | Use preset tones or custom clones |
| talkative people | Automatic detection and separation | Single-line output |
| Timeline editing | Detailed paragraph editing | No timeline concept |
| Input method | Audio/video file/platform link | Plain text input |
| Applicable Scenarios | Content localization and multilingual distribution | Audiobooks, navigation, customer service voice prompts |
Application scenarios of Dubbing v2
-
Podcast localizationThe program can be simultaneously translated and dubbed into 29 languages, covering the global market without the need for re-recording.
-
Cross-border e-commerce advertisingQuickly generate multilingual versions of a single video clip, significantly reducing the production costs of advertising.
-
Online EducationThe course videos are translated in batches while retaining the original voice characteristics of the instructors, enhancing the learning immersion of non-native speakers.
-
Film and television content distributionIndependent creators or small studios can achieve multilingual distribution of film and television works at low cost.
-
Corporate TrainingInternal training video materials are available in multiple languages, unifying brand messaging and accelerating knowledge transfer to global teams.