Cinema-grade AI videos with precise dual-prompt control
Wan 3.0 Prime Reference To Video composes clips from prompts plus reference media on the fast Prime tier.
Billing is per counted second: output duration plus the combined duration of any reference videos. Rates are $0.0624/s at 480p, $0.124/s at 720p, and $0.249/s at 1080p. Image and audio references are not billed as duration. Audio on or off does not change the rate. The pre-submit figure is an estimate; final charge is settled after the run.
| Parameter | Required | Type | Default | Range / Options | Description |
|---|---|---|---|---|---|
| prompt* | Yes (*) | string | Example prompt | Up to 20,000 characters | Scene; refer to refs as Image 1, Video 1, Audio 1. |
| reference_images | Conditional | array | example image | up to 10 | Reference images for subject or scene consistency. |
| reference_videos | Conditional | array | empty | up to 5 | Reference videos (1-15s each, 15s total). |
| reference_audios | Conditional | array | empty | up to 5 | Reference audio (15s total). |
| resolution | No | string | 720p | 480p, 720p, 1080p | Output resolution tier. |
| aspect_ratio | No | string | 16:9 | 16:9, 9:16, 1:1, 4:3, 3:4, adaptive | Output aspect ratio. |
| duration | No | integer | 5 | 2-30 | Output length in seconds. |
| prompt_extend | No | boolean | true | true / false | Auto-expand prompt; off can shorten wait. |
| enable_audio | No | boolean | true | true / false | Include synchronized audio. |
| seed | No | integer | random | 0-2147483647 | Seed for reproducible results. |
Provide at least one of reference_images, reference_videos, or reference_audios.
Cinema-grade AI videos with precise dual-prompt control
Create structured cinematic clips with audio, scene links, and prompt accuracy
Wan 3.0 Prime Text To Video makes clips from prompts fast
Turn text prompts into high quality videos with Tencent Hunyuan Video.
Generate fast, high quality videos from text with Kling 2.5 Turbo.
Lifelike characters, realistic physics, and stunning effects.
Wan 3.0 Prime Reference To Video builds a new clip from a text prompt plus reference images, videos, or audio, keeping subjects and motion consistent across the shot. It is built for character continuity, branded products, and multimodal storytelling on the fast Prime tier.
Supply up to 10 reference images, 5 reference videos, or 5 reference audio clips (with duration limits), then refer to them in the prompt as Image 1, Video 1, or Audio 1 in order. At least one reference type is required.
Wan 3.0 Prime Reference To Video uses the Prime model (wan3.0-video-prime) for faster inference while targeting the same Wan 3.0 visual tier. The multimodal reference workflow matches standard Wan 3.0 reference-to-video.
Output resolution is 480p, 720p, or 1080p; duration is 2 to 30 seconds; aspect ratios include 16:9, 9:16, and adaptive. Reference video and audio totals are capped—check the RunComfy parameter panel for exact limits.
Yes. Prototype in the RunComfy model UI, then call Wan 3.0 Prime Reference To Video through the RunComfy HTTP API with the same reference and prompt parameters.
Generations cost $0.0624 per counted second at 480p, $0.124 at 720p, and $0.249 at 1080p. Counted seconds are the finished clip plus the combined duration of any reference videos you attach (images and audio are not billed as duration). Because reference clips are measured after the run, the figure shown before you submit is an estimate. Audio does not change the rate. Trial credits are typically available for new users.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





