logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

Wan 3.0 Prime Reference To Video: Fast Multimodal Video on Models and API | RunComfy

wan-ai/wan-3.0-prime/reference-to-video

Wan 3.0 Prime Reference To Video combines a prompt with image, video, and audio references to build coherent clips with consistent subjects, motion, and sound on a faster Wan 3.0 tier.

Describe the scene, subject, motion, camera movement, lighting, and style. Refer to references as Image 1, Video 1, Audio 1 in order.
Image 1
Up to 10 reference images for subject, object, or scene consistency. At least one reference (image, video, or audio) is required.
Up to 5 reference videos (MP4 or MOV, 1-15s each, total no more than 15s) for motion or scene guidance.
Up to 5 reference audio clips (total no more than 15s) to guide sound or timing.
Output resolution tier. Use 480p for quick drafts and 1080p for higher-quality output.
Output aspect ratio of the generated video.
Length of the generated video in seconds. Range is 2-30. Start short to iterate, then increase once the motion looks right.
When on, the model may rewrite your prompt for richer scene detail before generation. Turn off for faster turnaround if your prompt is already precise.
When on, the output video includes a synchronized audio track. Turn off for a silent clip.
Random seed for reproducible results. Range is 0 to 2147483647.
Idle
The rate per counted second (generated video duration plus the combined duration of reference videos) is $0.0624 for 480p, $0.124 for 720p, and $0.249 for 1080p. Image and audio references are not billed as duration.

Introduction To Wan 3.0 Prime Reference To Video

Wan 3.0 Prime Reference To Video assembles coherent clips from a prompt plus image, video, or audio references on the fast Wan 3.0 Prime tier, with 2–30 second duration and aspect-ratio control. Trading fragile compositing for reference-guided consistency, it helps studios, ad teams, and creators keep characters, objects, motion, and sound on-brand across a shot. For developers, Wan 3.0 Prime Reference To Video on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: Character Consistency | Branded Product Scenes | Multimodal Storytelling

Wan-AI / Wan 3.0 Prime Reference To Video#


Wan 3.0 Prime Reference To Video composes clips from prompts plus reference media on the fast Prime tier.


Pricing#


Billing is per counted second: output duration plus the combined duration of any reference videos. Rates are $0.0624/s at 480p, $0.124/s at 720p, and $0.249/s at 1080p. Image and audio references are not billed as duration. Audio on or off does not change the rate. The pre-submit figure is an estimate; final charge is settled after the run.


Parameters#


ParameterRequiredTypeDefaultRange / OptionsDescription
prompt*Yes (*)stringExample promptUp to 20,000 charactersScene; refer to refs as Image 1, Video 1, Audio 1.
reference_imagesConditionalarrayexample imageup to 10Reference images for subject or scene consistency.
reference_videosConditionalarrayemptyup to 5Reference videos (1-15s each, 15s total).
reference_audiosConditionalarrayemptyup to 5Reference audio (15s total).
resolutionNostring720p480p, 720p, 1080pOutput resolution tier.
aspect_ratioNostring16:916:9, 9:16, 1:1, 4:3, 3:4, adaptiveOutput aspect ratio.
durationNointeger52-30Output length in seconds.
prompt_extendNobooleantruetrue / falseAuto-expand prompt; off can shorten wait.
enable_audioNobooleantruetrue / falseInclude synchronized audio.
seedNointegerrandom0-2147483647Seed for reproducible results.

Provide at least one of reference_images, reference_videos, or reference_audios.

Related Models

wan-2-1/fusionx/image-to-video

Cinema-grade AI videos with precise dual-prompt control

veo-3-1/first-last-frame-to-video

Create structured cinematic clips with audio, scene links, and prompt accuracy

wan-3.0-prime/text-to-video

Wan 3.0 Prime Text To Video makes clips from prompts fast

hunyuan/text-to-video

Turn text prompts into high quality videos with Tencent Hunyuan Video.

kling-2-5/turbo/text-to-video

Generate fast, high quality videos from text with Kling 2.5 Turbo.

luma-ray-2/image-to-video

Lifelike characters, realistic physics, and stunning effects.

Frequently Asked Questions

What is Wan 3.0 Prime Reference To Video used for?

Wan 3.0 Prime Reference To Video builds a new clip from a text prompt plus reference images, videos, or audio, keeping subjects and motion consistent across the shot. It is built for character continuity, branded products, and multimodal storytelling on the fast Prime tier.

How do references work in Wan 3.0 Prime Reference To Video?

Supply up to 10 reference images, 5 reference videos, or 5 reference audio clips (with duration limits), then refer to them in the prompt as Image 1, Video 1, or Audio 1 in order. At least one reference type is required.

How is Wan 3.0 Prime Reference To Video different from standard Wan 3.0 reference-to-video?

Wan 3.0 Prime Reference To Video uses the Prime model (wan3.0-video-prime) for faster inference while targeting the same Wan 3.0 visual tier. The multimodal reference workflow matches standard Wan 3.0 reference-to-video.

What output limits apply to Wan 3.0 Prime Reference To Video?

Output resolution is 480p, 720p, or 1080p; duration is 2 to 30 seconds; aspect ratios include 16:9, 9:16, and adaptive. Reference video and audio totals are capped—check the RunComfy parameter panel for exact limits.

Can developers use Wan 3.0 Prime Reference To Video through the RunComfy API?

Yes. Prototype in the RunComfy model UI, then call Wan 3.0 Prime Reference To Video through the RunComfy HTTP API with the same reference and prompt parameters.

How much does Wan 3.0 Prime Reference To Video cost on RunComfy?

Generations cost $0.0624 per counted second at 480p, $0.124 at 720p, and $0.249 at 1080p. Counted seconds are the finished clip plus the combined duration of any reference videos you attach (images and audio are not billed as duration). Because reference clips are measured after the run, the figure shown before you submit is an estimate. Audio does not change the rate. Trial credits are typically available for new users.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Seedance 2.5 Reference to Video 1080p
  • Seedance 2.5 1080p Text to video
  • Seedance 2.5 1080p
  • MiniMax H3 Open
  • Wan 2.6 Flash
  • Happy Horse 1.1 reference to video
  • View All Models →
Image Models
  • Qwen Image 3.0 Edit
  • Qwen Image 3.0 Pro Edit
  • Qwen Image 3.0
  • seedream 4.0
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of Wan 3.0 Prime Reference To Video

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...