Turns static visuals into cinematic motion with synced audio and natural camera flow
Wan 3.0 is Wan-AI's all-in-one video generation model that turns text, images, video, audio, files, and web pages into coherent clips with synchronized sound. This page provides Wan 3.0 Image to Video: give it a first-frame image and a prompt, and it animates that frame into a 2–30-second clip, with an optional last frame to pin how the shot ends. The same underlying model also runs text-to-video and reference-to-video, so a look you develop here transfers across the family.
| Wan 3.0 advantage | What it means for you |
|---|---|
| One model, many inputs | Wan 3.0 handles text, first/last-frame images, reference images, video, audio, documents, and links, so you rarely need to switch tools between ideas. |
| First-and-last-frame control | On this page you can set both the opening and closing frame, giving Wan 3.0 tighter control over how a shot begins and resolves than a first-frame-only animation. |
| Flexible 2–30 second output | Wan 3.0 renders short drafts or longer single takes, so quick iterations and finished shots come from the same endpoint. |
| Synchronized audio, optional | Every Wan 3.0 result can include a matching audio track, or you can export a silent clip—toggling audio does not change the price. |
The first table lists the controls exposed by the Wan 3.0 Image to Video tool on this page.
| Parameter | Required | Type | Default | Range / Options | How to choose |
|---|---|---|---|---|---|
prompt* | Yes (*) | String | Example prompt | Up to 20,000 characters | Describe the subject, its motion, the camera move, lighting, and style. Clear structure matters more than length. |
image_url* | Yes (*) | String | Example image | Image URL or Base64 | The first frame Wan 3.0 animates from. Use a sharp, well-lit, uncluttered image for the best subject preservation. |
end_image_url | No | String | empty | Image URL or Base64 | An optional last frame; set it when the ending pose or composition is fixed. |
resolution | No | String | 720p | 480p, 720p, 1080p | Use 480p for quick drafts and 1080p for higher-detail delivery. |
aspect_ratio | No | String | adaptive | adaptive, 16:9, 9:16, 1:1, 4:3, 3:4 | Leave adaptive to follow the input image, or force a ratio for a specific placement. |
duration | No | Integer | 5 | 2–30 seconds | Keep it short while refining motion, then raise it once the direction is confirmed. |
thinking_mode | No | Boolean | false | true / false | Enable for complex prompts that combine several movements or composition rules. |
enable_audio | No | Boolean | true | true / false | Keep on for a synchronized audio track; turn off for a silent clip. |
seed | No | Integer | Random | 0–2147483647 | Fix a seed to reproduce a result while you make small, comparable edits. |
The following table summarizes the wider Wan 3.0 model. Reference, document, and text-only inputs are available through the model's other modes and sibling tools, not as controls on this Image-to-Video page.
| Core dimension | Wan 3.0 |
|---|---|
| Model | wan3.0-video |
| Generation modes | Text-to-Video; Image-to-Video (first frame, or first + last frame); Reference-to-Video from images, video, and audio; plus document and web-link references. |
| Output duration | 2–30 seconds. With video input, input plus output stays within 30 seconds. A smart-duration option lets Wan 3.0 recommend a length. |
| Resolution | 480P, 720P, or 1080P. |
| Output aspect ratio | adaptive (recommended from the input) or 16:9, 4:3, 1:1, 3:4, 9:16. |
| Output audio | Optional synchronized audio, generated with the picture and on by default. |
| First / last-frame input | Up to one first_frame and one last_frame. Images: JPEG, JPG, PNG, BMP, or WEBP; 240–8000 pixels per side; aspect ratio up to 8:1; up to 20 MB each. |
| Reference images | Up to 10 reference images for subject, object, or scene consistency. |
| Reference videos | Up to 5 clips; MP4 or MOV; 1–15 seconds each; combined video up to 15 seconds; 240–4096 pixels per side; up to 100 MB per clip. |
| Reference audio | Up to 5 clips; WAV or MP3; combined audio up to 15 seconds; up to 15 MB each. |
| Document / web input | One document (DOCX, PDF, PPTX, XLSX, TXT, and similar, up to 50 pages) or one public web link. |
| Mode exclusivity | First/last-frame inputs cannot be combined with reference, document, or link inputs in the same request. |
| Prompt limit | Up to 20,000 characters. |
Wan 3.0 pricing depends on the resolution you pick and the generated duration. Enabling or disabling Wan 3.0 audio does not change the rate.
| Resolution | Price per second | 5s | 10s | 30s |
|---|---|---|---|---|
| 480p | $0.06 | $0.30 | $0.60 | $1.80 |
| 720p | $0.12 | $0.60 | $1.20 | $3.60 |
| 1080p | $0.25 | $1.25 | $2.50 | $7.50 |
For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.
Use this reusable Wan 3.0 structure:
[subject + defining details] + [one motion over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [audio] + [ending frame or constraint]
Use this Wan 3.0 comparison as a model-family guide; maximum resolution and duration may not be available together, and different tools can expose different inputs and controls. Based on publicly available information.
| Model | Resolution | Max duration | Audio | Standout |
|---|---|---|---|---|
| Wan 3.0 | 480p–1080p | 30s | Optional synchronized audio | One all-in-one model spanning text, first/last-frame, and reference-to-video from images, video, and audio, plus document and web inputs. |
| Kling 3.0 | Up to 4K | 15s | Native audio with lip-sync | Coordinates multi-shot camera changes and speaker-assigned dialogue for scripted ads and character scenes. |
| Hailuo 02 | 1080p | ~10s | None | Physics-focused motion and strong prompt adherence for silent action and product shots. |
| Seedance 2.5 | 1080p | 30s | Joint audio and video | Reference-heavy generation with synchronized speech and effects for identity-sensitive edits. |
| Veo 3.1 | 4K | 8s | Native dialogue and effects | First/last-frame and reference controls suited to cinematic transitions and assembled sequences. |
What sets Wan 3.0 apart is breadth: a single model that animates a first frame, honors a last frame, and can pull from images, video, audio, documents, and links—while generating optional synchronized sound. Choose Wan 3.0 Image to Video when you have an opening frame to bring to life, and move to the text or reference tools when your starting point changes.
Turns static visuals into cinematic motion with synced audio and natural camera flow
Animate between two images with smooth keyframe transitions using Pikaframes.
Turn text into detailed cinematic scenes with Dreamina 3.0 precision.
Open-weights image-to-video with optional last frame and native stereo audio.
Create synchronized prompt-based motion clips with precise audio and LoRA style control.
Seedance 2.5 Reference 480p: Multi-reference draft video at lower cost
Wan 3.0 image-to-video takes a single first-frame image and animates it into a short, coherent clip based on your text prompt. It is a good fit for turning product shots, portraits, lifestyle photos, or concept art into motion without rigging or manual keyframes.
You always provide a first-frame image, and you can optionally add a last-frame image when the ending pose or composition matters. Wan 3.0 then generates continuous motion that starts at your first frame and resolves toward the last one, giving you tighter control over how the shot begins and ends.
Because the first frame is used directly as the opening of the video, Wan 3.0 tends to keep the subject's look and framing consistent through the clip. Using a sharp, well-lit, uncluttered input image gives the best subject preservation.
Wan 3.0 supports 480p, 720p, and 1080p output, with 720p as a common default. Duration is adjustable from 2 to 30 seconds, and aspect ratio can be widescreen, vertical, square, or classic, or set to adaptive so it follows the input image. Check the RunComfy parameter panel for the exact current limits.
Yes. Wan 3.0 can produce a synchronized audio track along with the clip, and you can turn audio off when you only need silent footage. Toggling audio does not change the generation cost.
Input images should be common formats such as JPEG, PNG, or WEBP, with a reasonable resolution and aspect ratio. Very low-quality or heavily cluttered frames can reduce motion quality, so prefer a clean, high-resolution first frame. Limits may vary by provider settings.
Yes. You can prototype Wan 3.0 in the RunComfy model UI, then call the same model through the RunComfy HTTP API with identical parameters for production or automation. That lets you move from a browser test to an integrated pipeline without changing the model.
Generations consume usd or credits based on the output resolution and duration: $0.06 per second at 480p, $0.12 per second at 720p, and $0.25 per second at 1080p. New users typically get a free trial amount to try Wan 3.0 before committing to larger runs.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





