Seedance 2.0 Complete Prompting Guide
Learn how to write effective prompts for Seedance 2.0 text-to-video, image-to-video, and reference-to-video workflows.
Seedance 2.0 is a multimodal model for creative video generation. It can turn text, images, video, and audio references into clips with coherent motion, camera language, and audio-visual atmosphere. This guide explains how to choose the right generation mode and express a creative idea through clear, actionable prompts.
1. What Seedance 2.0 Can Do
1.1 Text-to-Video
Text-to-video begins with text alone. Describe the subject, action, environment, style, camera, and sound, and the model turns the idea into a complete moving scene.
This mode offers the most creative freedom. It is useful for concept previews, advertising ideas, atmospheric shorts, imagined scenes, and quick exploration when no visual assets are available. Without an image or video anchor, prompt specificity directly affects subject design, spatial relationships, and motion.
1.2 Image-to-Video
Image-to-video animates a still photograph, illustration, product image, or concept artwork. Because the image already defines appearance, structure, color, and the basic setting, the prompt should not merely repeat what is visible. Describe what happens next and how each layer moves.
You can control body movement, micro-expressions, hair and clothing, environmental effects such as wind, fog, and light, and camera behavior such as a push-in, pan, or locked shot. For illustrations, watercolor, animation, or 3D artwork, explicitly preserve the original medium so the result does not drift into another visual style.
1.3 Reference-to-Video
Reference-to-video is designed for work that needs stronger control and consistency. Images, video, and audio can each control a specific part of the result: character appearance, product details, action, camera movement, art direction, rhythm, or emotional development.
The key is to assign every asset a role. For example, @image1 can define appearance, @video1 can provide only action and camera rhythm, and @audio1 can drive emotional progression. The more references you use, the more clearly you must define priority to avoid conflicts in appearance, movement, and style.
1.4 Choose the Right Generation Mode
- You have only an idea and text description: choose text-to-video and build the visual world from scratch.
- You already have a strong frame, portrait, product image, or artwork: choose image-to-video and focus on how the existing image should move.
- You need to reuse appearance, action, camera language, style, or audio rhythm: choose reference-to-video and define the exact job of every asset.
2. Foundations of an Effective Prompt
Use these principles when writing Seedance 2.0 prompts:
- Be specific: describe visible and audible details such as appearance, action, setting, lighting, and sound. Replace vague words such as “beautiful” or “premium” with concrete details that explain the intended result.
- Organize logically: a reliable order is subject, action, environment, style, camera, and sound.
A useful prompt structure looks like this:
| Element | Description | Example |
|---|---|---|
| Subject / character | Main subject, appearance, clothing, or key object | A young violinist in a dark green coat carrying a wooden violin |
| Action / motion | What the subject does, including speed, direction, and rhythm | Walks through the station and begins playing beneath the clock |
| Environment / scene | Location, time, weather, depth, and background activity | An empty old railway station at dawn with mist drifting across the tracks |
| Style / atmosphere | Visual medium, color, lighting, and emotional tone | Cinematic, quiet, and hopeful, with cool blue shadows and soft golden morning light |
| Camera / framing | Shot size, angle, camera motion, focus, and composition | Begin wide, then slowly move into a medium shot |
| Audio / sound | Ambience, action sounds, music, or dialogue | Distant train ambience, soft wind, clear solo violin, no dialogue |
Combined into one complete prompt:
A young violinist in a dark green coat, carrying a weathered wooden violin, walks slowly through an empty old railway station at dawn, then stops and begins to play beneath the clock. Light fog drifts across the tracks. Cinematic realism, quiet and hopeful, with cool blue shadows and soft golden morning light. Begin with a wide shot, then make a slow dolly-in to a medium close-up with shallow depth of field. Distant train ambience, soft wind, clear solo violin, no dialogue.- Use precise keywords: define lighting, art direction, or camera language with phrases such as
soft lighting,watercolor style, orslow tracking shot. Avoid overlapping or contradictory terms. - Write important exclusions: reserve negative instructions for key restrictions, such as “do not change the character's appearance” or “do not make the image photorealistic.”
- Assign reference roles clearly: state whether each uploaded image, video, or audio file controls appearance, motion, camera behavior, or music.
- Test and iterate: the first result may not be perfect. Adjust only a few words, details, or priorities at a time so you can identify which instruction improved the output.
3. Seedance 2.0 Prompt Frameworks
3.1 Text-to-Video
Text-to-video is the most direct Seedance 2.0 workflow. Without image or video anchors, prompt structure and specificity matter even more.
Prompt
A solitary meteorologist in a bright orange weather suit stands on a black volcanic ridge as a vast thunderstorm approaches across the ocean. She raises a handheld sensor into the wind; her coat and loose straps whip violently while sheets of rain sweep across the rocks. A distant lightning strike illuminates the cloud layers, and she turns toward the flash with a focused expression. Photorealistic cinematic drama, cold steel-blue palette, wet reflective textures, strong backlight through the rain. Start with a wide establishing shot, then track slowly around her to a low-angle medium shot as the lightning flashes. Deep wind, rolling thunder, rain striking fabric and stone, no dialogue.- Begin with the subject: place the most important person, product, or object at the start so the model establishes a visual focus quickly.
- Describe motion clearly: action is the core of a video prompt.
- Define the rhythm: when pacing matters, use wording such as
a slow, meditative sequenceora fast-paced, high-energy montage. - Define the light: describe its source, direction, and texture, such as
soft, diffused morning light,harsh neon backlighting, orflickering candlelight casting warm shadows. - Add atmosphere: terms such as
tension-filled,whimsical,melancholic, oreuphorichelp establish emotional tone.
3.2 Image-to-Video
When you provide a still photograph, artwork, or concept image, the prompt should tell Seedance 2.0 how to bring the existing frame to life.
Unlike text-to-video, you do not need to rebuild the scene in words. Focus on what happens next and how each layer moves.
Design Motion in Layers
Prompt
Animate with fine red dust sweeping across the cracked ground in the foreground. The astronaut's loose fabric straps flutter gently in a steady wind, and faint condensation gathers along the inside edge of the visor. In the background, red warning lights flicker on the abandoned outpost, and two reconnaissance drones circle lazily above the antenna towers. Apply a slow, contemplative camera push-in toward the astronaut. The mood is lonely and monumental, with a muted rust-and-cyan color grade and soft, diffused sunset lighting.
A strong image-to-video prompt treats the scene as multiple layers that can move independently.
A common beginner mistake is describing only the subject while ignoring environmental changes. Coordinating foreground, subject, and background usually produces more natural, cinematic results.
- Foreground: elements nearest the lens, such as drifting leaves, flickering candles, or reflections on water.
- Midground: the main subject and core action, such as a slow turn or a horse shifting its weight.
- Background: depth-building elements, such as moving clouds, distant flags, or people walking far behind the subject.
The three layers do not need equal intensity. Make the main action clear, then support it with subtler foreground and background movement.
Use Micro-Movements to Convey Emotion
When a face is visible, small changes in expression and posture can make the subject feel alive without introducing unstable movement.
Small eye movements, a faint squint, visible breathing, or a collar lifted by wind often conveys more emotion than an exaggerated performance.
Prompt
Animate with a barely perceptible shift in the fisherman's gaze—his eyes slowly tracking something distant on the horizon. A faint squint tightens around his eyes. His jacket collar flutters softly. Waves reflect subtly in his eyes. The camera remains completely static, locked off. The atmosphere is deeply contemplative and nostalgic, desaturated with warm tones with soft coastal light.
Preserve Abstract and Artistic Styles
Image-to-video can animate drawings, illustrations, and concept art as well as photographs.
For these references, explicitly preserve the original brushwork, edges, color, texture, and medium. This prevents the animation from becoming over-rendered or visually inconsistent.
Prompt
Animate this scene while fully preserving the dreamy watercolor storybook aesthetic—soft, luminous color washes, delicate painterly textures, glowing gold details, and diffused starlight throughout. The little boy gently plays the harp, his fingers softly plucking the strings as they shimmer and vibrate with golden light. White birds circle gracefully around the harp, while swallows glide and flutter across the starry sky in smooth, flowing paths. The animation should feel hand-crafted, poetic, and delicate, never sharp or digital. A gentle, whimsical atmosphere with celestial blue, soft white, and warm golden tones.
Image-to-Video Keyword Reference
These are useful building blocks, not fixed formulas.
| Category | Practical keywords |
|---|---|
| Motion | gently, barely perceptible, slowly drifting, rhythmically swaying, subtly rippling |
| Atmosphere | mist rolling in, particles of dust, heat haze, soft bokeh, volumetric light rays |
| Character life | micro-expression shift, eyes slowly tracking, breath visible, hair softly lifted by wind |
| Camera | locked off, slow push-in, subtle drift, gentle handheld sway, rack focus |
| Style retention | maintain painterly texture, preserve film grain, honor the original color palette |
3.3 Reference-to-Video
Preserve and Transform
A reference-to-video prompt can be divided into two explicit sections: Preserve and Transform. This tells Seedance 2.0 what must remain consistent and what should change instead of making the model guess.
Prompt
(Preserve) Retain all original movement, choreography, timing, and body posture of the dancer exactly as they appear in the source video. Maintain the original camera angle and framing throughout.
(Transform) Re-stylize the entire visual environment as an ethereal, otherworldly forest glade. Replace the studio floor with a carpet of luminous, floating flower petals. Surround the dancer with slow-moving fireflies and drifting luminescent spores. The dancer's costume should transform into a flowing, translucent gown that trails light. Apply a dreamlike, fantasy aesthetic with soft teal and lavender tones, volumetric god rays filtering through ancient trees. Film grain texture, cinematic quality.Style Transfer: Define a New Visual Language
When you want to preserve source-video content but change the entire style, do more than name a broad aesthetic. Describe the target palette, lighting, materials, image texture, and period details.
Also state which source elements must remain, such as main actions, performance order, dialogue timing, or camera movement.
Prompt
Preserve all dialogue timing, gestures, and the camera position. Re-stylize the entire scene in the visual language of a Studio Ghibli animated feature—soft, hand-drawn cel animation aesthetic, warm and richly textured backgrounds, characters rendered with expressive Ghibli-style proportions. The café transforms into a charming, vintage European bakery with afternoon sunlight streaming through lace curtains. Palette is warm, creamy, and inviting. Gentle ambient sounds of clinking cups and soft piano music implied in the visual atmosphere.Multimodal Fusion
Multimodal generation combines images, video, and audio in one request. This freedom also adds complexity: each asset can introduce a different style, rhythm, palette, or emotional tone. The goal is to establish one creative direction and prevent references from competing.
Establish Creative Priority
Think of the input assets as departments in a film production. One defines visual identity, another shapes action and camera behavior, and another controls emotional rhythm. Give each asset one clear job.
Prompt
@image1 is the primary visual authority—the protagonist’s face, hairstyle, blue hair streak, cybernetic eye implant, black tactical clothing, illuminated cyan details, boots, and messenger bag must remain exactly consistent throughout the entire video. Do not redesign, replace, or simplify any part of her appearance.
@video1 serves exclusively as the movement, parkour choreography, body-mechanics, and camera reference. Apply its exact sprinting rhythm, barrier vault, wall run, landing impact, low slide, and final acceleration to the protagonist from @image1. Preserve the original sequence, timing, spatial direction, tracking shots, camera orbit, and low floor-level camera movement, but do not carry over the gray training outfit, stunt performer’s identity, warehouse, or any other visual element from @video1.
@audio1 sets the emotional rhythm and editing intensity of the entire sequence. During the restrained opening pulses, begin with controlled running and a smooth low-angle tracking shot. As the percussion builds, increase the protagonist’s speed, environmental motion, and camera energy. Synchronize the vault, wall push, slide, and strongest camera movements with the major rhythmic accents. At the musical drop, reveal the chase at full intensity and finish on the final impact.
Fusion: Equal-Weight References
When two or more assets should contribute equally to a new visual world, explain how their qualities combine rather than naming one as the dominant reference.
Prompt
Fuse the visual identities of @image1 and @image2 equally into a single, cohesive world—a retro-futurist city that exists at the intersection of 1930s art deco grandeur and contemporary neon Tokyo nightlife. Neither should dominate; the architecture carries the geometric elegance of @image2 while glowing with the saturated neon palette and wet-reflective streets of @image1. Animate a slow, gliding aerial camera drift through this world, unhurried and contemplative. Let @audio1 dictate the pace entirely—every camera movement should feel as languid and swinging as the jazz rhythm. The atmosphere is nostalgic, mysterious, and quietly beautiful.

Use Audio as the Primary Driver
Music and sound design can control a video's structure from beginning to end. Describe how the scene responds: restraint during quiet passages, increasing environmental motion as the score grows, and the strongest visual change at the crescendo.
Prompt
Let @audio1 be the architect of this entire video. Begin in near-silence: a static, locked-off shot of the lighthouse from @image1—still, barely animated, only the faintest movement of stormy clouds. As the orchestral score begins to swell, incrementally increase the intensity of the environment—waves grow larger, lightning begins to flash in the distance, the wind picks up, the lighthouse beam begins to rotate. By the time the score reaches its full crescendo, the scene should be a breathtaking storm in full fury—crashing waves, torrential rain, dramatic lightning strikes illuminating the cliff face, the lighthouse beam cutting through the chaos. The visuals and music must feel inseparable, as if one created the other. Cinematic, photorealistic, deeply dramatic.
4. Summary
A good prompt is not a pile of adjectives. It assigns creative functions clearly: who or what the subject is, what happens, how the environment responds, how the camera observes the scene, and how sound shapes emotional development.
Text-to-video builds a scene from scratch. Image-to-video brings an existing frame to life. Reference-to-video locks appearance, movement, style, and rhythm through clearly assigned asset roles.
Start with a clear idea, then adjust action, camera behavior, and atmosphere in small steps. Clear intent is more reliable than trying to specify every detail at once.
Seedance 2.5 Prompting Guide (Part 2): Video Editing, Extension, and Final Assembly
Learn Seedance 2.5 prompts for instruction-based editing, reference-image editing, dialogue localization, extension, automatic assembly, and seamless transitions.
MiniMax H3 Prompting Guide
Master MiniMax H3 video prompting through prompt structure, reference-asset roles, camera, sound, complete examples, and generated results.