Unified AI model for refined scene editing, style match, and smooth video refits
LTX 2.5 Fast Audio To Video builds a clip whose picture timing follows a supplied audio bed—dialogue, music, or SFX—in a speed-optimized mode. Upload audio (2–20 seconds), optionally add a start image and a prompt, and generate 1080p video aligned to the track. Use the text-to-video or image-to-video Fast pages when you are not driving the cut from audio.
| Advantage | What it means for you |
|---|---|
| Audio-led timing | Picture length and phrasing follow the uploaded clip—useful for music spots, VO shorts, and track-locked ads. |
| Optional visual anchor | Add a start image to lock identity or product look; omit it and rely on the prompt alone. |
| Fast previews | Speed-focused mode for iterating cut ideas before a heavier final render elsewhere. |
| Simple delivery controls | Aspect ratio auto / 16:9 / 9:16, plus guidance scale to tighten prompt adherence. |
Controls on this LTX 2.5 Fast Audio To Video page:
| Parameter | Required | Type | Default | Range / Options | How to choose |
|---|---|---|---|---|---|
audio_url* | Yes (*) | Audio URL | Example clip | MP3, OGG, WAV, M4A, AAC; 2–20 seconds | Use a clean bed; output length follows this clip. |
image_url | No | Image URL | — | JPG, JPEG, PNG, WEBP, GIF, AVIF | Anchor identity or product; leave empty for prompt-only look. |
prompt | Conditional | String | Example prompt | Up to 5,000 characters | Required when no image is provided; describe motion timed to the audio. |
guidance_scale | No | Number | 5 | 1–50 | Higher values follow the prompt more tightly (defaults near 9 when an image is present on the provider side). |
aspect_ratio | No | String | auto | auto, 16:9, 9:16 | auto follows the image when present; otherwise defaults toward 16:9. |
LTX 2.5 Fast Audio To Video is billed per second of input audio at 1080p:
| Input audio | Price at $0.14/s |
|---|---|
| 5s | $0.70 |
| 10s | $1.40 |
| 15s | $2.10 |
| 20s | $2.80 |
For batches, total ≈ audio_seconds × $0.14 × output count.
Tips for LTX 2.5 Fast Audio To Video:
guidance_scale only if the prompt is being ignored.Improved prompt example
> A vocalist in a black turtleneck sings into a vintage ribbon microphone in a dim booth, headphones on. Soft side key light, slow push-in matching the vocal phrasing. Clean studio realism, no text, no watermark.
Unified AI model for refined scene editing, style match, and smooth video refits
Streamline video refinements with seamless scene continuity for creators.
Create lifelike video motion fast with Seedance Pro for design pros
Turn text prompts into high quality videos with Tencent Hunyuan Video.
Animate stills into native 4K cinematic clips with start-end frame guidance and synchronized sound.
Animate a still photo into smooth 720P or 1080P video from one prompt.
LTX 2.5 Fast Audio To Video generates video timed to a supplied audio clip. It fits music-driven shorts, dialogue teasers, VO spots, and ads keyed to a track when picture must follow sound.
An audio URL is required (2–20 seconds; MP3, WAV, M4A, AAC, or OGG). An optional start image can lock identity, and a prompt describes motion; the prompt is required if you do not provide an image.
Generation is billed and delivered at 1080p on this page. Clip length follows the input audio duration rather than a separate duration control. Aspect ratio can be auto, 16:9, or 9:16.
Guidance scale (1–50, default 5) controls how tightly the result follows the prompt. Raise it gradually if the text is being ignored; very high values can look over-constrained.
Yes. Add an optional image_url as the first frame to stabilize face, product, or packaging identity while the audio drives timing. Without an image, rely on a clear prompt.
Yes. Configure a run in the RunComfy Web UI, then call the same LTX 2.5 Fast Audio To Video model via the HTTP API with identical parameters for automation.
At 1080p, billing is $0.14 per second of input audio. For example, 10 seconds of audio costs $1.40 per output. Generations consume USD/credits; new users typically receive a free trial amount.
Pick LTX 2.5 Fast Audio To Video when a finished music bed, VO, or dialogue take must set the cut timing. Use text-to-video Fast when you are still inventing sound and picture together from a prompt.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





