logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

Wan 3.0: Image-to-Video, First-Last Frame, Audio on Models and API | RunComfy

wan-ai/wan-3.0/image-to-video

Wan 3.0 animates a first-frame image into a coherent video with optional last-frame control, flexible duration, aspect ratio, and audio for polished short clips.

Describe the motion, action, camera movement, lighting, and style you want the first frame to animate into.
First-frame image the model animates into video. Use a clear, high-quality image for better subject preservation.
Optional last-frame image used to guide the ending pose, composition, or scene of the video.
Output resolution tier. Use 480p for quick drafts and 1080p for higher-quality output.
Output aspect ratio. Use adaptive to match the input image proportions.
Length of the generated video in seconds. Range is 2-30. Start short to iterate, then increase once the motion looks right.
Enable deeper prompt interpretation for complex scenes with multiple movement or composition requirements.
When on, the output video includes a synchronized audio track. Turn off for a silent clip.
Random seed for reproducible results. Range is 0 to 2147483647.
Idle
The rate is $0.06 per second for 480p, $0.12 per second for 720p, and $0.25 per second for 1080p.

Introduction To Wan 3.0

Wan 3.0 image-to-video brings a single first-frame image to life as a cinematic clip, with optional last-frame guidance, 2-30 second duration, aspect-ratio control, and synchronized audio. Trading slow manual keyframing for prompt-driven motion that keeps your subject and scene consistent, it suits product marketers, portrait and lifestyle creators, and social teams shipping short clips fast. For developers, Wan 3.0 on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: Product & Marketing Clips | Portrait Animation | Vertical Social Video

Why Choose Wan 3.0#


Wan 3.0 is Wan-AI's all-in-one video generation model that turns text, images, video, audio, files, and web pages into coherent clips with synchronized sound. This page provides Wan 3.0 Image to Video: give it a first-frame image and a prompt, and it animates that frame into a 2–30-second clip, with an optional last frame to pin how the shot ends. The same underlying model also runs text-to-video and reference-to-video, so a look you develop here transfers across the family.


Wan 3.0 advantageWhat it means for you
One model, many inputsWan 3.0 handles text, first/last-frame images, reference images, video, audio, documents, and links, so you rarely need to switch tools between ideas.
First-and-last-frame controlOn this page you can set both the opening and closing frame, giving Wan 3.0 tighter control over how a shot begins and resolves than a first-frame-only animation.
Flexible 2–30 second outputWan 3.0 renders short drafts or longer single takes, so quick iterations and finished shots come from the same endpoint.
Synchronized audio, optionalEvery Wan 3.0 result can include a matching audio track, or you can export a silent clip—toggling audio does not change the price.

Best Use Cases#


  • Product and marketing videos: Use Wan 3.0 to turn a product or packshot into a short promotional clip with controlled motion and framing.
  • Portrait and lifestyle animation: Let Wan 3.0 animate a portrait, fashion, or lifestyle still into a natural, on-brand moment.
  • Social media content: Produce vertical, square, or widescreen Wan 3.0 clips from a single image for Reels, Shorts, Stories, or feed placements.
  • Creative prototyping: Test camera direction, motion, and visual style from a fixed frame with Wan 3.0 before committing to a full production pass.

How It Works#


  1. Set the opening frame: Upload a clean, high-resolution first-frame image that Wan 3.0 will bring to life.
  2. Describe the motion: Tell Wan 3.0 what moves, how the camera moves, and how the light and scene develop over time.
  3. Guide the ending (optional): Add a last-frame image when the final pose, composition, or scene matters, and Wan 3.0 will bridge the two.
  4. Choose delivery format: Pick a resolution, an aspect ratio, a 2–30-second duration, and whether to generate audio, then review the Wan 3.0 result and refine one instruction at a time.

Parameters#


The first table lists the controls exposed by the Wan 3.0 Image to Video tool on this page.


ParameterRequiredTypeDefaultRange / OptionsHow to choose
prompt*Yes (*)StringExample promptUp to 20,000 charactersDescribe the subject, its motion, the camera move, lighting, and style. Clear structure matters more than length.
image_url*Yes (*)StringExample imageImage URL or Base64The first frame Wan 3.0 animates from. Use a sharp, well-lit, uncluttered image for the best subject preservation.
end_image_urlNoStringemptyImage URL or Base64An optional last frame; set it when the ending pose or composition is fixed.
resolutionNoString720p480p, 720p, 1080pUse 480p for quick drafts and 1080p for higher-detail delivery.
aspect_ratioNoStringadaptiveadaptive, 16:9, 9:16, 1:1, 4:3, 3:4Leave adaptive to follow the input image, or force a ratio for a specific placement.
durationNoInteger52–30 secondsKeep it short while refining motion, then raise it once the direction is confirmed.
thinking_modeNoBooleanfalsetrue / falseEnable for complex prompts that combine several movements or composition rules.
enable_audioNoBooleantruetrue / falseKeep on for a synchronized audio track; turn off for a silent clip.
seedNoIntegerRandom0–2147483647Fix a seed to reproduce a result while you make small, comparable edits.

  • Required field.

The following table summarizes the wider Wan 3.0 model. Reference, document, and text-only inputs are available through the model's other modes and sibling tools, not as controls on this Image-to-Video page.


Core dimensionWan 3.0
Modelwan3.0-video
Generation modesText-to-Video; Image-to-Video (first frame, or first + last frame); Reference-to-Video from images, video, and audio; plus document and web-link references.
Output duration2–30 seconds. With video input, input plus output stays within 30 seconds. A smart-duration option lets Wan 3.0 recommend a length.
Resolution480P, 720P, or 1080P.
Output aspect ratioadaptive (recommended from the input) or 16:9, 4:3, 1:1, 3:4, 9:16.
Output audioOptional synchronized audio, generated with the picture and on by default.
First / last-frame inputUp to one first_frame and one last_frame. Images: JPEG, JPG, PNG, BMP, or WEBP; 240–8000 pixels per side; aspect ratio up to 8:1; up to 20 MB each.
Reference imagesUp to 10 reference images for subject, object, or scene consistency.
Reference videosUp to 5 clips; MP4 or MOV; 1–15 seconds each; combined video up to 15 seconds; 240–4096 pixels per side; up to 100 MB per clip.
Reference audioUp to 5 clips; WAV or MP3; combined audio up to 15 seconds; up to 15 MB each.
Document / web inputOne document (DOCX, PDF, PPTX, XLSX, TXT, and similar, up to 50 pages) or one public web link.
Mode exclusivityFirst/last-frame inputs cannot be combined with reference, document, or link inputs in the same request.
Prompt limitUp to 20,000 characters.

Pricing#


Wan 3.0 pricing depends on the resolution you pick and the generated duration. Enabling or disabling Wan 3.0 audio does not change the rate.


ResolutionPrice per second5s10s30s
480p$0.06$0.30$0.60$1.80
720p$0.12$0.60$1.20$3.60
1080p$0.25$1.25$2.50$7.50

For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.


Prompting & Reference Tips#


Use this reusable Wan 3.0 structure:


[subject + defining details] + [one motion over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [audio] + [ending frame or constraint]


  • Lead with the subject: Name what the first frame contains, then describe how Wan 3.0 should move it.
  • Separate camera from subject: Describe subject motion and camera motion independently for cleaner Wan 3.0 results.
  • Use a sharp first frame: Blur, clutter, or low resolution in the input carries into the video.
  • Anchor the ending: Reach for a last-frame image when a specific final pose or composition is non-negotiable.
  • Turn on thinking mode: Enable it when a Wan 3.0 shot layers several movements or composition rules.
  • Fix the seed: Lock a seed once a take looks right so you can compare small prompt edits fairly.

How Wan 3.0 Compares#


Use this Wan 3.0 comparison as a model-family guide; maximum resolution and duration may not be available together, and different tools can expose different inputs and controls. Based on publicly available information.


ModelResolutionMax durationAudioStandout
Wan 3.0480p–1080p30sOptional synchronized audioOne all-in-one model spanning text, first/last-frame, and reference-to-video from images, video, and audio, plus document and web inputs.
Kling 3.0Up to 4K15sNative audio with lip-syncCoordinates multi-shot camera changes and speaker-assigned dialogue for scripted ads and character scenes.
Hailuo 021080p~10sNonePhysics-focused motion and strong prompt adherence for silent action and product shots.
Seedance 2.51080p30sJoint audio and videoReference-heavy generation with synchronized speech and effects for identity-sensitive edits.
Veo 3.14K8sNative dialogue and effectsFirst/last-frame and reference controls suited to cinematic transitions and assembled sequences.

What sets Wan 3.0 apart is breadth: a single model that animates a first frame, honors a last frame, and can pull from images, video, audio, documents, and links—while generating optional synchronized sound. Choose Wan 3.0 Image to Video when you have an opening frame to bring to life, and move to the text or reference tools when your starting point changes.


More Models to Try#


  • Wan 3.0 Text to Video: Generate a clip from a prompt alone when you have no starting image.
  • Wan 3.0 Reference to Video: Combine images, video, and audio references for character and object consistency.
  • Seedance 2.5 Text to Video: Produce fast, low-cost draft renders from text.
  • Kling 3.0: Animate a start image with an optional end frame and synchronized audio.

Official Resources#


  • Wan 3.0 video generation API reference (Alibaba Cloud Model Studio)
  • Wan model family overview

Related Models

kling-2-6/pro/image-to-video

Turns static visuals into cinematic motion with synced audio and natural camera flow

pikaframes

Animate between two images with smooth keyframe transitions using Pikaframes.

dreamina-3-0/pro/text-to-video

Turn text into detailed cinematic scenes with Dreamina 3.0 precision.

minimax-h3-open/image-to-video

Open-weights image-to-video with optional last frame and native stereo audio.

ltx-2-19b/text-to-video/lora

Create synchronized prompt-based motion clips with precise audio and LoRA style control.

seedance-2.5/reference-to-video/480p

Seedance 2.5 Reference 480p: Multi-reference draft video at lower cost

Frequently Asked Questions

What is Wan 3.0 used for in image-to-video workflows?

Wan 3.0 image-to-video takes a single first-frame image and animates it into a short, coherent clip based on your text prompt. It is a good fit for turning product shots, portraits, lifestyle photos, or concept art into motion without rigging or manual keyframes.

How does the first-and-last-frame option in Wan 3.0 work?

You always provide a first-frame image, and you can optionally add a last-frame image when the ending pose or composition matters. Wan 3.0 then generates continuous motion that starts at your first frame and resolves toward the last one, giving you tighter control over how the shot begins and ends.

How well does Wan 3.0 preserve the subject from the input image?

Because the first frame is used directly as the opening of the video, Wan 3.0 tends to keep the subject's look and framing consistent through the clip. Using a sharp, well-lit, uncluttered input image gives the best subject preservation.

What resolutions, durations, and aspect ratios does Wan 3.0 image-to-video support?

Wan 3.0 supports 480p, 720p, and 1080p output, with 720p as a common default. Duration is adjustable from 2 to 30 seconds, and aspect ratio can be widescreen, vertical, square, or classic, or set to adaptive so it follows the input image. Check the RunComfy parameter panel for the exact current limits.

Can Wan 3.0 generate audio together with the video?

Yes. Wan 3.0 can produce a synchronized audio track along with the clip, and you can turn audio off when you only need silent footage. Toggling audio does not change the generation cost.

What input limits should I know before using Wan 3.0?

Input images should be common formats such as JPEG, PNG, or WEBP, with a reasonable resolution and aspect ratio. Very low-quality or heavily cluttered frames can reduce motion quality, so prefer a clean, high-resolution first frame. Limits may vary by provider settings.

Can developers use Wan 3.0 through the RunComfy API?

Yes. You can prototype Wan 3.0 in the RunComfy model UI, then call the same model through the RunComfy HTTP API with identical parameters for production or automation. That lets you move from a browser test to an integrated pipeline without changing the model.

How much does it cost to generate with Wan 3.0 on RunComfy?

Generations consume usd or credits based on the output resolution and duration: $0.06 per second at 480p, $0.12 per second at 720p, and $0.25 per second at 1080p. New users typically get a free trial amount to try Wan 3.0 before committing to larger runs.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • MiniMax H3 Open
  • FLUX 3 Image to Video
  • MiniMax H3 Open Image to Video
  • Wan 2.6 Flash
  • Happy Horse 1.1 reference to video
  • Seedance 1.5 Pro Text to Video
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2
  • Qwen Image Edit 2511 LoRA
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of Wan 3.0

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...