Multimodal AI video model with native audio for text, image, and reference inputs.
| Parameter | Required | Type | Default | Range / Options | Description |
|---|---|---|---|---|---|
| prompt* | Yes (*) | string | A lone cyclist crosses a rain-swept bridge at blue hour as the camera tracks alongside, reflected city lights shimmering below, cinematic lighting, natural motion. | Free input | Describe the scene, subject, action, camera movement, lighting, and motion. |
| aspect_ratio | No | string | 16:9 | 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9 | Frame shape for the generated video. |
| resolution | No | string | 720p | 360p, 540p, 720p, 1080p | Output quality tier shown as video resolution. |
| duration | No | integer | 5 | 1-15 | Requested output duration in seconds, from 1 through 15. |
| generate_audio_switch | No | boolean | false | Free input | Generate an optional native audio track with the video. |
| generate_multi_clip_switch | No | boolean | false | Free input | Enable a multi-shot sequence within the generated video. |
| seed | No | integer | 0 | Free input | Seed used to reproduce or vary a generation. |
[Limited Time Offer] Half price. Audio and 1080p are included at no extra cost. Billing is per generated second.
| Resolution | Price |
|---|---|
| 360p | $0.012/s |
| 540p | $0.017/s |
| 720p | $0.022/s |
| 1080p | $0.022/s |
Turning audio on does not change the rate. After this Limited Time Offer ends, regular rates start at $0.044/s for 720p and $0.089/s for 1080p.
Multimodal AI video model with native audio for text, image, and reference inputs.
Transforms static characters into smooth motion clips for flexible creative workflows
Create lifelike scenes with synced audio and visual fidelity.
Transforms input clips into synced animated characters with precise motion replication.
Efficient video transformation with cinematic motion and design precision.
Cinematic 4K reference-to-video at $0.419 per second of output.
PixVerse V6 Text To Video is configured for text to video generation with prompt direction, output controls, and an optional audio track. Outputs can run from 1 to 15 seconds at 360p, 540p, 720p, or 1080p.
PixVerse V6 Text To Video fits filmmakers, technical artists, designers, agencies, and product teams creating shot concepts, campaign assets, product motion, or social video. It is especially useful when prompt-level direction and repeatable parameters matter.
Give PixVerse V6 Text To Video separate instructions for subject action and camera movement, then add framing, lighting, atmosphere, and pacing. Keep the requested action realistic for the selected 1-15 second duration.
PixVerse V6 Text To Video includes a generate_audio_switch option that is off by default. Turning it on requests native audio and does not change the per-second rate.
PixVerse V6 Text To Video exposes 360p, 540p, 720p, and 1080p, with an integer duration from 1 through 15 seconds. Available aspect-ratio control depends on the task form.
PixVerse V6 Text To Video includes a seed parameter for controlled iteration. The same seed and inputs can improve repeatability, although generated-video systems may still show some variation.
Yes. Teams can test PixVerse V6 Text To Video in the RunComfy model UI and then send the same exposed parameters through the HTTP API. This supports moving a validated request into an application or automation workflow.
PixVerse V6 Text To Video is billed per generated second. Limited Time Offer rates are $0.012/s at 360p, $0.017/s at 540p, and $0.022/s at 720p or 1080p, with or without audio. Your total depends on duration.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





