Turn static visuals into smooth motion with Hailuo 2.3 for rapid, realistic video creation.
Seedance 2.5 Reference to Video 480p uses your reference material to drive generation at 480p. Blend up to 9 images, 3 videos, and 3 audio files into a single guided generation: images steer identity and style, videos carry camera motion and rhythm, and audio sets the mood — all combined via one text prompt, with optional synchronized native audio.
| Advantage | What it means for you |
|---|---|
| Multi-reference control | Combine up to 9 images, 3 videos, and 3 audio files so several sources guide one coherent result. |
| Stronger consistency | Reference images plus a clear prompt help anchor identity, wardrobe, and tone across frames. |
| Native audio in one pass | Generate synchronized speech, effects, and music with the clip, or turn audio off for silent video. |
| Up to 30-second clips | Direct a longer single shot with steadier quality than stitched short takes. |
The table below lists the controls exposed by the Seedance 2.5 Reference to Video 480p tool on this page.
| Parameter | Required | Type | Default | Range / Options | How to choose |
|---|---|---|---|---|---|
prompt* | Yes (*) | String | Example prompt | Chinese ~≤500 characters or English ~≤1000 words recommended | Describe the action and camera; references anchor identity, motion, and mood. |
images | No | Array (image URLs) | Example image | up to 9 | jpeg, png, webp, bmp, tiff, gif; steer identity and style. |
videos | No | Array (video URLs) | [] | up to 3 | mp4, mov; ~2–15 s each; carry camera motion and rhythm. |
audios | No | Array (audio URLs) | [] | up to 3 | wav, mp3; ~2–15 s, under 15 MB; set the mood. |
aspect_ratio | No | String | 16:9 | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive | Match the destination frame; adaptive lets the model pick the closest ratio. |
duration | No | Integer | 5 | 4–30 seconds, in 1-second steps | Short clips for a single action; longer only when the prompt has a clear arc. |
generate_audio | No | Boolean | true | true / false | Leave on for synchronized speech, effects, and music; turn off for silent video. |
\* Required field. Only the prompt is required; references are optional but recommended for consistent results.
The rate is $0.14 per second without a reference video, and $0.085 per second with a reference video.
Improved prompt example
> Using the greenhouse reference for environment and style, animate a calm cinematic walk: slow push-in down the central aisle as leaves tremble softly, dust motes drift through warm side light, glass panels softly fogging, shallow depth of field, continuous smooth motion, no text, no watermark.
Turn static visuals into smooth motion with Hailuo 2.3 for rapid, realistic video creation.
Transform still visuals into cinematic motion clips with smooth, realistic transitions and creative flexibility.
Generate cinematic clips faster with multimodal references, lip-sync, and camera control
Extend an audio track at the start, end, or both with matching style
Millisecond lipsync, emotion-aware realism, and flexible video design.
Animate a single image into a smooth video with Kling 2.1 Standard.
It guides a 480p clip with reference images and optional video or audio, keeping identity, wardrobe, and style consistent. It fits consistent-character videos, product-reference clips, and style-locked brand videos.
You attach reference material — up to 9 images, 3 short videos, and 3 audio clips — and describe the shot in the prompt. References anchor identity, wardrobe, style, motion, and sound, while the prompt guides action and camera. Only the prompt is strictly required.
Up to 9 reference images, 3 reference videos, and 3 reference audio files (reference videos and audio about 2–15 seconds each; audio under 15 MB). Duration is a whole number of seconds from 4 to 30 (default 5). Aspect ratio can be 16:9 (default), 9:16, 1:1, 4:3, 3:4, 21:9, or adaptive. Output is fixed at 480p.
No. It can run from a text prompt plus images alone. Add short reference videos or audio when you want stronger motion or mood guidance.
Yes. generate_audio is on by default, so the model can output synchronized speech, effects, and music. Turn it off when you only need silent video.
Yes. Prototype in the RunComfy model UI, then call the same template through the API with matching fields (prompt, images, videos, audios, aspect_ratio, duration, generate_audio). Generations consume credits on both paths.
Without a reference video, total price = output video duration × $0.14. With a reference video, total price = (input reference video duration + output video duration) × $0.085. Example: 10s reference video + 10s output = 20s × $0.085 = $1.70. Image and audio references are not billed.
Both share the same reference-guided path and 4–30-second window; this page outputs at 480p. Draft here, then reuse the winning references on the 720p page for finals.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





