Create fast, audio-enhanced visuals from text prompts
Controls on this LTX 2.5 Pro Audio To Video page:
| Parameter | Required | Type | Default | Range / Options | Description |
|---|---|---|---|---|---|
audio_url* | Yes (*) | Audio URL | Example clip | MP3, OGG, WAV, M4A, AAC; 2–20 seconds | Driving bed; output timing follows this clip. |
image_url | No | Image URL | — | JPG, JPEG, PNG, WEBP, GIF, AVIF | Optional first frame to lock identity or product look. |
prompt | Conditional | String | Example prompt | Up to 5,000 characters | Required when no image is provided; describe motion timed to the audio. |
guidance_scale | No | Number | 9 | 1–50 | Higher values follow the prompt more tightly. |
aspect_ratio | No | String | auto | auto, 16:9, 9:16 | auto follows the image when present. |
LTX 2.5 Pro Audio To Video is billed per second of input audio at 1080p:
| Input audio | Price at $0.19/s |
|---|---|
| 5s | $0.95 |
| 8s | $1.52 |
| 10s | $1.90 |
| 15s | $2.85 |
| 20s | $3.80 |
For batches, total ≈ audio_seconds × $0.19 × output count.
Create fast, audio-enhanced visuals from text prompts
Smart editing tool for refined video transfers and motion-based scene adjustments.
Generate videos from text prompts with audio using Wan 2.5 Preview.
Generate video between start and end frames with optional audio
Create lifelike avatars via multimodal synthesis with Omnihuman 1.5.
Reshape a source clip from a text prompt with native audio.
LTX 2.5 Pro Audio To Video generates 1080p video timed to a supplied 2–20 second audio clip in a quality-focused mode. It fits music spots, dialogue teasers, and soundtrack-led ads that need picture locked to sound.
Both follow an uploaded audio bed, but LTX 2.5 Pro Audio To Video targets delivery-oriented fidelity at the Pro rate. Use Fast when you want cheaper track-timed previews before committing to a final.
Audio is required (2–20 seconds). You may add an optional start image and a prompt; if no image is provided, the prompt is required. Aspect ratio and guidance scale are also available.
Choose auto, 16:9, or 9:16. Auto follows the start image when present, or defaults toward 16:9 without an image. Check the RunComfy parameter panel for the current defaults.
Yes. Provide a start image to anchor faces, products, or packaging, and use the prompt to describe motion timed to the audio. Without an image, the prompt alone drives the look.
Yes. Prototype in the RunComfy Web UI, then call the same LTX 2.5 Pro Audio To Video model via the HTTP API with matching parameters for automation and production workflows.
Billing is per second of input audio at 1080p: $0.19 per second. Generations consume USD/credits on RunComfy; new users typically receive a free trial amount.
Common formats such as MP3, WAV, M4A, AAC, and OGG are supported, with clip length between 2 and 20 seconds. Output timing follows the uploaded audio length.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





