Seedance 2.0 Tutorial - A Complete User Manual and Prompt Guide for AI Video Creation
Seedance 2.0 marks a significant step in AI video creation, moving from 'prompt-based lottery' to 'director-level precise control.' Through multimodal input and @reference mechanisms, creators can manage video content with the same precision as a real film crew, using images to define style and video content...
Seedance 2.0 marksAIVideo creation from "Prompt wordsLottery draws have entered an era of "director-level precision control," with models employing...MultimodalWith input and @quote mechanisms, creators can direct their work like a real film crew, using images to define style, videos to determine camera movement, audio to set the rhythm, and text to define the plot, completely moving beyond simply "writing."Prompt wordsThis tutorial systematically breaks down the model into eleven core capabilities and five practical templates to help users overcome the passive situation of "leaving it to fate."fastMaster this new creative language and truly unleash your potential in commercial-grade video production.AIHis potential as a director.
What is Seedance 2.0?
Seedance 2.0 is a product launched by ByteDance.MultimodalAIVideo creation platform. The platform's biggest breakthrough lies in its transformation.AIThe interactive way of video creation—from the old "write-to-get" approachPrompt wordsThe "leave it to fate" model has been upgraded to allow creators to precisely control every creative element, much like a true director. Users simultaneously input four different types of materials—images, videos, audio, and text—clearly specifying the function of each material, which is then organically integrated to generate a complete video work.
Before you begin creating, you need to understand the various input limitations and output characteristics that Seedance 2.0 can accept.
- Image inputThe system supports up to 9 images being uploaded.
- Video inputThe system supports up to 3 video files uploaded.
- Audio inputThe system supports MP3 audio files, with a maximum of 3 files that can be uploaded and a total duration not exceeding 15 seconds.
- Text inputThe system accepts natural language descriptions, both Chinese and English.
- File total limit: The maximum number of files that can be uploaded in total is 12.
- Generation time: The final generated video can be freely selected between 4 and 15 seconds, and can be flexibly adjusted according to actual needs.
- Sound output. The video generated by Seedance 2.0 willautomaticIt comes with sound effects and background music, requiring no additional processing.
practicalRecommendation: More materials are not necessarily better. Prioritize uploading materials that have the greatest impact on the visual style or rhythm, and allocate the 12 file slots reasonably to avoid wasting them on secondary content.
In most cases, the universal reference is sufficient, as it supports various reference inputs.up to dateThe way Seedance 2.0 can be maximized.
Step 1: Choose the right entry point
After opening the JiDream platform and finding Seedance 2.0, you will see two different entry points.
- First and last frame entryUse this when uploading only the first frame image and a text description.
- All-in-one reference portal. needMultimodalUse this when combining (images, videos, audio, and text).
How to choose? Remember one principle:If the source material consists of only one image and text, use the first and last frames; if the source material consists of more than one image, or involves video or audio, use the all-around reference.
Step 2: Upload materials
Click the upload button and select a file from your local drive. Images, videos, and audio files can all be dragged and dropped directly into the upload window. Once uploaded successfully, all materials will appear in the input area; hovering the mouse over them will allow you to preview the content.
Small suggestionBefore uploading, decide which materials are most crucial. You can only upload 12 files in total; prioritize uploading materials that have the greatest impact on the visual style and pacing.
Step 3: Assign tasks to each material using "@" (most crucial)
This step is the core operation of Seedance 2.0, and it's something that many beginners easily overlook.
After uploading the materials, you need toPrompt wordsThe `@materialName` tag tells the model what each material is for. The model won't guess; if it's not clearly defined, it might use it incorrectly.
How to invoke @:
- Method 1Typing an "@" character directly into the input box will...automaticA list of uploaded materials will pop up. Click on the material you want to use, and it will appear in the input box.
- Method 2Clicking the "@" button in the parameter toolbar next to the input box will also bring up the material list.
Example of the correct way to write @:
- Specify first frame and reference@Image1 serves as the first frame, referencing the camerawork of @Video1; @Audio1 is used for background music.
- Designated character imageThe girl in picture 1 is the main character, and the boy in picture 2 is the supporting character.
- Designated camera movement reference: All camera movements and transitions were completely referenced from @Video 1
- Specified scenario referenceThe scene on the left is referenced in image 3, and the scene on the right is referenced in image 4.
- Specified action referenceThe character reference in @Image 1 and the dance moves in @Video 1
- Specified timbre reference: Narrator's voice reference @Video 1
Tips for avoiding pitfallsWhen you have a lot of source material, always double-check that each @ reference is correct. Using an image as a video reference, or mistaking character A's icon for character B, will result in a very messy generated model. Hover your mouse over the source material you @'d to preview it and avoid insertion errors.
Step 4: Write it downPrompt words
After assigning tasks, all that's left is to describe the visuals and actions you want using natural language.
WritePrompt wordsFour tips:
Tip 1Write in segments according to timeline. If there are multiple shots or plot twists in the video, it is recommended to describe them in segments by the second.
- 0-3 seconds of footageThe male protagonist holds up a basketball, looks up at the camera, and says, "I just wanted a drink, am I about to time travel...?"
- 4-8 seconds of footageSuddenly, the camera shakes violently, and the scene changes to a rainy night in an old house, where a female protagonist dressed in ancient costume looks coldly in the direction of the camera.
- 9-13 seconds of footageThe camera cuts to a person dressed in Ming Dynasty clothing...
This approach to model writing allows for a more accurate grasp of the rhythm and content of each scene.
Technique 2Please clarify whether it's "reference" or "editing." These are different concepts. "Reference the camera movement of @Video 1" means borrowing its camera movement style to generate new content; "Replace the girl in @Video 1 with a female opera performer" means making modifications to the original video. Clearly specifying this is crucial for the model to respond correctly.
Tip 3Describe the camera language in detail. Don't be afraid to write a lot; the model's understanding is very strong now. It recognizes all the technical terms like push, pull, pan, tilt, tracking shot, circle, overhead shot, low-angle shot, one-shot, Hitchcock zoom, fisheye lens... If it doesn't understand the technical terms, that's okay too; you can describe it in plain language, such as "the camera slowly turns from behind to the front."
Tip 4: Add transition descriptions to continuous actions. If you want the character to perform a series of smooth actions, remember to write the transition relationship, such as "the character transitions directly from jumping to rolling, keeping the action smooth and continuous", to avoid unnatural jump cuts in the screen.
Seedance 2.0's Ten Core Capabilities
Capability 1: Significantly improved basic image quality
Seedance 2.0 has undergone a comprehensive upgrade at the underlying level, resulting in more reasonable physics, smoother movements, and a more stable style. Seedance 2.0 represents a qualitative leap in its core image generation capabilities:
- The laws of physics are more reasonableThe movement of clothes, splashing water, and collisions are all more realistic.
- More natural and fluid movementsThe characters' walking, running, and complex movements are no longer stiff.
- More accurate command understandingYou said "a girl elegantly hangs out to dry her clothes," does it really understand what "elegant" means?
- Maintain a more stable styleThe visuals remain consistent throughout, without any sudden changes in style.
Ability two:MultimodalFree combination
This is the core upgrade of Seedance 2.0 – the ability to use any material as a “reference”.
formula: Seedance 2.0 = MultimodalReference (can reference everything) + Strong creative generation + Precise instruction understanding
Things that can be referenced include:
- Actions, special effects, form
- Camera movement and cinematic language
- Character design and scene style
- Sound, musical rhythm
practicalSkillHow to write:Prompt words
- I have the first frame image, and I'd also like to refer to the video animation."@Figure 1 is the first frame, referencing the fighting action in @Video 1."
- Extend existing videos"Extend @Video1 by 5 seconds" (Also select 5 seconds for the generated duration).
- Integrating multiple videos"Add a scene between @Video1 and @Video2, with the content xxx".
- Use the sound from the videoNo need to upload audio separately; you can refer to the video directly.
- Continuous action"The character transitions directly from jumping to rolling, maintaining a smooth and fluid movement."
Capability 3: Comprehensive Improvement in Consistency
Seedance 2.0 has put a lot of effort into this aspect. After uploading a character reference image, the appearance, clothing, and posture of the character remain consistent throughout the entire video. The same applies to product displays; a bag can be shown from multiple angles and rotated, and the material details from the front and sides are not lost.
Elements that can maintain consistency:
- Facial features (facial features, skin tone, expression style)
- Clothing details (texture, color, pattern)
- Brand elements (logo, font, color scheme)
- Scene style (lighting, atmosphere, color tone)
Capability 4: Precise replication of camera movement and actions
It only takes two steps: upload a video of your favorite camera movement as a reference, and write "referencing the camera movement effect of @Video 1".
The model can recognize camera movements in reference videos (push-pull, pan, tilt, circle, tracking shot, zoom, one-shot, etc.) and apply the same camera movement logic to new content.
Types of camera movements that can be replicated:
- Hitchcock Zoom
- Surround shooting
- One-shot
- Push, pull, rock, move
- Low-angle shot
- Aerial view
Capability 5: Accurate replication of creative templates and special effects
See a cool ad idea, transition effect, or movie clip? Upload it directly as a reference. The model can recognize the movement rhythm, visual structure, and camera language, and replicate it to create your own version.
Replicable creative types:
- Creative transitions (puzzle breaking, particle dissipation, pupil passing through, etc.)
- Advertising film style
- MV rhythm editing
- Movie special effects shots
- Transformation/Face Swap Effects
Skill 6: Video Extension and Seamless Connection
Already have a satisfactory video and want to continue filming? Or want to add back a scene? The video extension function can handle it all.
- Extend backwardUpload an existing video and write "Extend @Video1 by X seconds" along with a description of the new footage.
- Extend forwardWrite "Extend forward by X seconds" followed by a description of the preceding events.
Usage Rules:
- Tell the model "Extend @Video1 by X seconds".
- The generation duration should be selected as the duration of the extended portion (for example, if the extension is 5 seconds, the generation length should be 5 seconds).
- New storylines and visual descriptions can be added to the extended section.
- Supports forward or backward extension
Capability 7: More realistic sound
Videos generated by Seedance 2.0 come with built-in sound effects and background music, and the sound quality is much better than before.
Several sound-related gameplay options:
- Reference toneUpload a video or audio clip and let the model imitate the speaking tone or narration style.
- Multilingual dialogueThe character can speak multiple languages, including Chinese, English, Spanish, and Korean, and their emotional expression is quite effective.
- Multi-character dialogueIt supports multiple characters speaking their own lines in a single video. Successful examples include cat and dog stand-up comedy, dialogues from period dramas, and tactical conversations in military themes.
- Dialect supportSomeone successfully got a character to order milk tea in Sichuan dialect, and the effect was quite authentic.
- Sound effect matchingThe system can accurately generate environmental sound effects such as footsteps, thunder, crowd noise, and equipment collisions.
Skill 8: One-shot filming for greater continuity
Seedance 2.0 has made significant progress in this area. Upload multiple images of different scenes and describe it as "a single, continuous tracking shot, following a runner up stairs, through a corridor, onto a rooftop, and finally overlooking the city," and the model can achieve natural transitions between scenes without noticeable breaks. More complex single-shot sequences can also be implemented.
SkillMultiple images are arranged in sequence, and the model will present these scenes sequentially in a single, continuous shot.
Skill Nine: Video Editing Skills
You already have a video, but don't want to start from scratch and only want to modify a part of it? Now you can directly use an existing video as input to make targeted modifications.
- Role replacementReplace A in the video with B, keeping the actions and expressions the same. For example, "replace the female lead singer in video 1 with the male lead singer in picture 1, and completely imitate the actions of the original video."
- Plot subversionKeep the setting and characters the same, but completely rewrite the plot. Someone changed a touching video of moon-gazing on a bridge into a plot twist where the male lead pushes the female lead into the water. Another person changed a tense bar negotiation into a hilarious twist where the male lead pulls out a huge bag of snacks.
- Element modificationChange hairstyles, add props, change backgrounds. For example, "Change the woman's hairstyle in video 1 to long red hair, and in image 1, the great white shark slowly emerges with half its head behind her."
- Brand integrationInsert brand elements into existing videos. For example, add a close-up of a paper bag with the brand logo in a fried chicken video.
Skill 10: Music Timing
Upload a rhythmic music video as a reference, and the model can recognize the changes in the music's beat, allowing the scene transitions to precisely match the beat.
- Basic checkpointsUpload source images and music reference videos, and write "match the beat according to the rhythm of the scenes in the @video".
- Dynamic beatWrite: "The characters in the picture are more dynamic, the overall picture style is more dreamlike, the picture has strong tension, and the framing of the reference picture can be changed according to the needs of the music."
- Scenic Spots: Multiple landscape photos with music, with the caption "Refer to the rhythm of the scenes in @video, and match the style of the scenes and the rhythm of the music during transitions."
Skill Eleven: More Effective Emotional Performance
The characters' stiff facial expressions and awkward emotional transitions have always been a problem.AIThe video had the same old problem. Seedance 2.0 has made significant improvements in this regard.
Reference Video
- General writing style"Refer to this video."
- Better way to write"Refer to the camera movement and transition effects in @Video 1".
Use images
- General writing style"Use this image."
- Better way to write"@Image 1 is the first frame, and the character's appearance is referenced from @Image 2".
Rhythm control
- General writing style"Make a rhythmic video."
- Better way to write"Refer to the rhythm of the visuals and the timing of the music in @Video 1."
Extended video
- General writing style"Extend the video".
- Better writing"Extend @Video 1 by 5 seconds and add xxx content."
Replace character
- General writing style"Change to someone else."
- Better way to write"Replace the female protagonist in @Video 1 with the image of @Image 1, and completely imitate the actions of the original video."
The golden formula: @materials + usage instructions + specific scene descriptions + timeline (optional)
- Don't forget @I uploaded the materials, but...Prompt wordsWithout @ references, uploading is pointless. The model won't automatically guess the purpose of each image.
- @Don't mislabelProblems are most likely to occur when there is a lot of material. (After writing...)Prompt wordsThen, take 10 seconds to check that each @ reference is correct.
- Extend the video selection durationIf you want to extend the generation time by 5 seconds, choose the 5-second generation duration. Choosing a longer duration will generate more unnecessary content.
- The reference video should not be too long.The total duration is capped at 15 seconds, and shorter videos are more accurate. If you only want to reference a specific shot, extracting those key seconds is sufficient.
- Generate multiple times: AIThe generation process inherently involves randomness; the same input might yield vastly different results in three runs. Don't give up just because you're not satisfied with the first attempt; try several rounds and choose the best one.
- FirstSimplePost-complexIf you are a beginner, it is recommended to start with "one picture + text", and then add video and audio references as you become familiar with them, proceeding step by step.
The core value of Seedance 2.0 lies in...MultimodalInput and @reference mechanism willAIVideo creation from "Prompt wordsThe "blind box" concept has transformed into "director-level precision control"—creators can now manage content like a real film crew, using images to define style, videos to determine camera movement, audio to set the pacing, and text to define the plot, achieving truly controllable generation. Eleven capabilities already cover commercial scenarios such as e-commerce advertising, short drama trailers, and brand promotional videos. While there's still room for optimization in extremely complex narratives, "everyone can be a director" has moved from a slogan to a productive reality. We recommend saving this and getting started immediately: begin with a single image and gradually add more.MultimodalThrough repeated adjustments, I mastered this new creative language.