Generate cinematic 3-15s videos from text with optional sound.
FLUX 3 is Black Forest Labs' next unified model, trained jointly on images, video, and audio so motion and sound reason from the same foundation. This page focuses on its video side and previews what FLUX 3 Video will feel like in production.
Note: FLUX 3 Video is coming soon. This page previews the experience while Black Forest Labs finalizes the public model and API.
Why creators reach for the FLUX 3 Video experience:
A quick read on what FLUX 3 Video can do:
Black Forest Labs describes FLUX 3 Video as one multimodal model with several generation modes, and every output carries native audio. The officially confirmed tasks are:
Beyond those modes, FLUX 3 Video also delivers:
The reason FLUX 3 Video stands out is that picture and sound are designed together, not stitched on afterward. In a single pass it can produce dialogue, foley, ambience, and music timed to the action: a door slams with its own impact, an alarm pulses in step with a warning light, and a character can speak in the same shot the camera moves through.
That synchronization is a by-product of how the model was trained. Black Forest Labs reports that video prediction accounted for more than 95% of FLUX 3's training compute, which forced the model to learn contact, motion, weight, and cause and effect, in other words how the physical world actually behaves. Audio is then learned as a consequence of that world model, so sounds line up with the on-screen events that cause them and speech tracks lip movement. For creators, that means motion reads as physically plausible and the soundtrack reinforces the edit instead of fighting it.
These tips help you get predictable output while FLUX 3 Video is previewed:
Because FLUX 3 Video is still in early access, the clearest evidence comes from Black Forest Labs' own preliminary evaluations, 10-second 720p clips generated with audio. In those side-by-side tests, reviewers preferred FLUX 3 over a broad field: Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in up to 69%, Kling v3 Pro in 60%, and both Seedance 2.0 and Gemini Omni Flash in 52%. The numbers are early and expected to move, but they anchor the comparisons below in capability rather than hype.
FLUX 3 Video is a preview of a model still in limited early access, so it helps to set expectations:
Treat FLUX 3 Video as a look at where the model is heading: its capabilities and limits will keep shifting as it moves toward general availability.
Black Forest Labs is rolling FLUX 3 AI out as one multimodal foundation with several editions, each released in phases after its own early-access period. Based on the official launch plan:
Black Forest Labs has confirmed open weights are planned, and once the FLUX 3 Dev backbone is available, a community FLUX 3 LoRA workflow, including FLUX 3 dev LoRA fine-tuning for recurring characters, styles, and brand looks, is the natural next step, mirroring the large LoRA ecosystem that grew around earlier FLUX releases.
For most teams, FLUX 3 Video will handle day-to-day clips, while FLUX 3 Dev and FLUX 3 dev LoRA open the door to self-hosted customization.
In short, FLUX 3 Video is coming soon on RunComfy: one multimodal foundation that turns prompts and reference media into cinematic clips with synchronized native audio.
Generate cinematic 3-15s videos from text with optional sound.
Generate cinematic motion clips with precise control and audio sync
Animate stills into native 4K cinematic clips with start-end frame guidance and synchronized sound.
Turn text prompts into high quality videos with Tencent Hunyuan Video.
Turn still portraits into expressive, lifelike videos with control and precision.
Transforms reference clips into 1080p short videos with precise motion and voice alignment.
FLUX 3 Video is coming soon. This preview lets you explore the FLUX 3 Video experience while Black Forest Labs finalizes the unified multimodal model and its API. Until the FLUX 3 backend is switched on, the live input fields and limits reflect the model currently powering this preview.
FLUX 3 Video is designed to turn prompts and optional image, video, and audio references into short cinematic clips with native audio. It targets ad creative, film previsualization, and branded storytelling where reference consistency and synchronized sound matter.
FLUX 3 Video is positioned to bring Black Forest Labs' work into motion and audio under one multimodal foundation, with five officially confirmed tasks — text-to-video, image-to-video, video-to-video, video + audio continuation, and keyframe-to-video — plus native audio, multilingual dialogue, and clips up to 20 seconds in a single generation. Exact gains depend on your content, so compare clips on the same prompt once the model goes live.
FLUX 3 Video generates clips up to 20 seconds in a single pass, roughly double the 10-second limit common to most video models, which leaves room for dialogue beats and slower dramatic shots. Clips can also be chained into multi-shot sequences that run for minutes while characters stay consistent. On the model currently powering this preview, duration runs from 4 to 15 seconds; the full FLUX 3 Video model is designed to reach 20 seconds.
Yes. Native audio is a core goal for FLUX 3 Video, and on the model currently powering this preview you can enable audio generation to output synchronized speech, SFX, and music, which supports lip-synced, multilingual clips. You can disable it for silent video. Lip-sync quality depends on prompt clarity and the scene you describe.
FLUX 3 was trained with video prediction as the bulk of its compute, so the model had to learn contact, motion, weight, and cause and effect, essentially how the physical world behaves. Native audio is learned on top of that world model, which is why sounds line up with the on-screen events that cause them and speech tracks lip movement. In practice, FLUX 3 Video is noted for expressive faces and physically plausible motion.
Reference guidance is central to FLUX 3 Video. In practice, reference images plus a clear prompt help anchor identity, wardrobe, and tone across frames, and FLUX 3 is designed to carry a subject from a source clip into a new scene. On the model currently powering this preview you can attach reference images, videos, and audio within the supported limits to reinforce consistency.
On the model currently powering this preview, prompts are recommended at Chinese ~≤500 characters or English ~≤1000 words. You can attach up to 9 reference images, 3 reference videos (2–15 seconds each), and 3 reference audio clips (2–15 seconds, under 15 MB). Only the prompt is required. Check the live parameter panel, since limits will change once FLUX 3 Video goes live.
Resolution options are 480p, 720p (default), 1080p, and 4K. Aspect ratio can be adaptive (default, where the model picks the closest ratio) or fixed at 16:9, 9:16, 4:3, 3:4, 1:1, or 21:9. Duration is a whole number of seconds from 4 to 15 on the model currently powering this preview, with a default of 5; the full FLUX 3 Video model is designed to reach 20 seconds in a single generation.
FLUX 3 AI is planned as one foundation with several editions. Alongside hosted FLUX 3 Video, Black Forest Labs has announced FLUX 3 Image, FLUX 3 Action (the FLUX-mimic robotics work), and FLUX 3 Dev, an open-weight multimodal backbone. Black Forest Labs has confirmed open weights are planned; once the FLUX 3 Dev backbone lands, expect a community FLUX 3 LoRA workflow, including FLUX 3 dev LoRA training for recurring characters, styles, and brand looks.
Prototype in the RunComfy model UI, then call the same template through the RunComfy API using identical Input fields (prompt, images, videos, audios, aspect_ratio, duration, resolution, generate_audio, seed). Validate prompts and media limits in the UI first, then use your account API key and credits for automated jobs.
Generations consume credits based on the generated video duration on the model currently powering this preview, Seedance 2.0 Pro: $0.08 per second for 480p, $0.175 per second for 720p, $0.40 per second for 1080p, and $0.90 per second for 4K. See the pricing shown on this page, which will update when FLUX 3 Video launches.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





