logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

FLUX 3 Video: Multimodal Video Generation with Native Audio on Models and API | RunComfy

blackforestlabs/flux-3/video

FLUX 3 Video creates clips up to 20 seconds with synchronized native audio. It is currently in limited early access and is not yet live on RunComfy.

Text prompt for the video (Chinese ~≤500 characters, English ~≤1000 words recommended).
Reference images for multimodal reference mode (0–9). Support jpeg、png、webp、bmp、tiff、gif.
Reference videos for multimodal reference mode (0–3). Support mp4、mov. Video duration must be between 2 and 15 seconds.
Reference audio for multimodal reference mode (0–3). Support wav、mp3. Audio duration must be between 2 and 15 seconds. Size should be less than 15MB.
Default is adaptive (the model picks the closest ratio; the actual ratio is returned on task query).
Integer seconds in [4, 15].
When true, the model outputs video with synchronized audio (speech, SFX, music).
Random seed for the video generation.
When `web_search` is included, the model may run an online search depending on the prompt (e.g. specific products, current weather), which can improve factual freshness but increases latency.
Idle
FLUX 3 Video is coming soon. Now the model is Seedance 2.0 Pro. Billed on total video seconds (input + output). The rate per second is $0.08 for 480p, $0.175 for 720p, $0.40 for 1080p, and $0.90 for 4k.

Introduction To FLUX 3 Video Creation

FLUX 3 Video is coming soon on RunComfy. It is Black Forest Labs' unified multimodal model applied to video and audio, turning prompts and reference media into cinematic clips with synchronized native audio. Trading frame-by-frame editing, manual masking, and separate dubbing for prompt-driven generation, FLUX 3 Video is aimed at marketing teams, agencies, film previz artists, and game studios that need cinematic short clips fast. For developers, FLUX 3 Video on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: High-Conversion Video Ads | Shot-Accurate Film Previsualization | Multi-language Lip-Synced Brand Narratives

Black Forest Labs / FLUX 3 Video#


FLUX 3 is Black Forest Labs' next unified model, trained jointly on images, video, and audio so motion and sound reason from the same foundation. This page focuses on its video side and previews what FLUX 3 Video will feel like in production.


Note: FLUX 3 Video is coming soon. This page previews the experience while Black Forest Labs finalizes the public model and API.


FLUX 3 Video Highlights#


Why creators reach for the FLUX 3 Video experience:


  • Skip the edit suite: a written brief stands in for timeline editing, dubbing, and manual cleanup.
  • Direct it like a shot: in FLUX 3 Video, camera angle, mood, and pacing all come from the words you write.
  • Sound in the same pass: every clip arrives with synchronized native audio, not a separate dubbing step.
  • Fits real deliverables: results are shaped for ads, previz, and social placements, not just test footage.

FLUX 3 Video at a glance#


A quick read on what FLUX 3 Video can do:


  • Clip length: up to 20 seconds in a single generation, roughly double the 10-second ceiling common to most video models, with room to chain shots into multi-minute sequences.
  • Audio: native, synchronized sound generated together with the picture, and optional when you want silent footage.
  • Inputs: text prompts plus image, video, and audio references (up to about ten image references), so it edits and remixes as readily as it generates from scratch.
  • Resolution: 720p during early access, with higher resolutions expected as the model matures.
  • Aspect ratios: a wide spread from vertical 9:16 to ultrawide 21:9 for social, film, and product formats.
  • Styles: candid camcorder, animation, and polished cinematics from one model.
  • Status: limited early access, so FLUX 3 Video is coming soon on RunComfy.

FLUX 3 Video Key Capabilities#


Black Forest Labs describes FLUX 3 Video as one multimodal model with several generation modes, and every output carries native audio. The officially confirmed tasks are:


  • Text-to-video: generate a fully voiced clip up to 20 seconds from a written prompt alone.
  • Image-to-video: animate from a starting frame, or use supplied images as visual references to shape the motion.
  • Video-to-video: carry central elements of a source clip, the same character or object, into a new scene or context.
  • Video + audio continuation: extend an input video and its soundtrack instead of starting over.
  • Keyframe-to-video: define moments and let it fill the controlled transition between them.

Beyond those modes, FLUX 3 Video also delivers:


  • Multilingual dialogue with mouth movement that tracks the speech.
  • Agentic multi-shot chaining that links clips into sequences running for minutes while characters stay consistent.
  • A broad range of styles and aspect ratios, from candid camcorder footage to animation and cinematics.
  • Strong typography and animated design rendered directly in the frame.
  • Grounded realism: Black Forest Labs notes FLUX 3 Video is already standout at expressive faces and at matching sounds to on-screen physical events.

Native audio and world-grounded motion in FLUX 3 Video#


The reason FLUX 3 Video stands out is that picture and sound are designed together, not stitched on afterward. In a single pass it can produce dialogue, foley, ambience, and music timed to the action: a door slams with its own impact, an alarm pulses in step with a warning light, and a character can speak in the same shot the camera moves through.


That synchronization is a by-product of how the model was trained. Black Forest Labs reports that video prediction accounted for more than 95% of FLUX 3's training compute, which forced the model to learn contact, motion, weight, and cause and effect, in other words how the physical world actually behaves. Audio is then learned as a consequence of that world model, so sounds line up with the on-screen events that cause them and speech tracks lip movement. For creators, that means motion reads as physically plausible and the soundtrack reinforces the edit instead of fighting it.


FLUX 3 Video Prompt & Reference Tips#


These tips help you get predictable output while FLUX 3 Video is previewed:


  • Be specific about camera and motion (medium close-up, slow push-in, handheld vs locked-off).
  • Use images for what must stay stable (face, costume, logo); use text for what should evolve (action, mood).
  • Keep reference videos and audio within 2-15 seconds, and reference audio under 15 MB.
  • With audio on, name the dialogue tone or ambient sound you want to hear.
  • Fix the seed while iterating so changes come from the prompt, not randomness.
  • Simplify the prompt if the result feels noisy or contradictory.

How FLUX 3 Video compares to other models#


Because FLUX 3 Video is still in early access, the clearest evidence comes from Black Forest Labs' own preliminary evaluations, 10-second 720p clips generated with audio. In those side-by-side tests, reviewers preferred FLUX 3 over a broad field: Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in up to 69%, Kling v3 Pro in 60%, and both Seedance 2.0 and Gemini Omni Flash in 52%. The numbers are early and expected to move, but they anchor the comparisons below in capability rather than hype.


FLUX 3 Video vs FLUX.2#


  • Image-only to motion: every prior FLUX, including FLUX.2, generates stills only; FLUX 3 is the first to fold video and synchronized audio into the same foundation, so FLUX's reference handling, prompt adherence, and typography now carry into clips.
  • Shared DNA: the traits that made FLUX images dependable, holding a character, product, or logo steady, are what FLUX 3 Video extends across a shot.
  • New territory: motion, temporal stability, and on-beat audio are fresh problems for an image-first lab, which is exactly why the early-access window exists.

FLUX 3 Video vs Seedance 2.0#


  • The data: in Black Forest Labs' early tests, FLUX 3 was preferred over Seedance 2.0 in 52% of comparisons, essentially even with a mature, proven multimodal video model.
  • Capability: Seedance 2.0 mixes text, image, video, and audio in one request and returns roughly fifteen seconds of multi-shot 480p-4K video with native, synchronized multi-track audio plus edit and extend loops. FLUX 3 Video matches that multimodal ambition and, per Black Forest Labs, is particularly strong at expressive faces, matching sound to physical events, and multilingual dialogue.
  • How to choose: if you must ship today, Seedance 2.0 is proven; FLUX 3 Video is the one to watch as its public release approaches.

FLUX 3 Video vs Kling, Runway, and Luma#


  • The data: FLUX 3 was preferred over Kling v3 Pro in 60%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93% of Black Forest Labs' early comparisons.
  • What drives the gap: those margins track FLUX 3 Video's stated strengths, expressive human faces, audio locked to on-screen physical events, multilingual dialogue, and consistent multi-shot chaining, the details that make a short clip read as finished rather than as a test render.

What to expect from FLUX 3 Video in early access#


FLUX 3 Video is a preview of a model still in limited early access, so it helps to set expectations:


  • Strong today: long, coherent clips, dialogue-driven and dramatic scenes, and native synchronized audio all hold up well in early hands-on testing.
  • Still maturing: image-to-video reference adherence can be inconsistent in early builds, and very fast, high-energy action does not yet carry the kinetic intensity of the most motion-focused rivals.
  • Rolling out: resolution starts around 720p with higher resolutions expected, while pricing and open weights are still to be confirmed by Black Forest Labs.

Treat FLUX 3 Video as a look at where the model is heading: its capabilities and limits will keep shifting as it moves toward general availability.


FLUX 3 Video editions and the wider FLUX 3 AI family#


Black Forest Labs is rolling FLUX 3 AI out as one multimodal foundation with several editions, each released in phases after its own early-access period. Based on the official launch plan:


  • FLUX 3 Video: video and audio generation and editing through APIs and private weight access, the experience previewed on this page.
  • FLUX 3 Image: image synthesis and editing, also delivered through APIs and private weight access.
  • FLUX 3 Action (FLUX-mimic): action prediction for robotics, built with partners such as mimic robotics and already tested on production lines at Audi.
  • FLUX 3 Dev: an open-weight multimodal backbone spanning image, video, and audio generation plus action prediction, for teams that want to run and adapt the model themselves.

Black Forest Labs has confirmed open weights are planned, and once the FLUX 3 Dev backbone is available, a community FLUX 3 LoRA workflow, including FLUX 3 dev LoRA fine-tuning for recurring characters, styles, and brand looks, is the natural next step, mirroring the large LoRA ecosystem that grew around earlier FLUX releases.


For most teams, FLUX 3 Video will handle day-to-day clips, while FLUX 3 Dev and FLUX 3 dev LoRA open the door to self-hosted customization.


FLUX 3 Video Official Resources#


  • Black Forest Labs

In short, FLUX 3 Video is coming soon on RunComfy: one multimodal foundation that turns prompts and reference media into cinematic clips with synchronized native audio.

Related Models

kling-video-o3/standard/text-to-video

Generate cinematic 3-15s videos from text with optional sound.

veo-3-1/text-to-video

Generate cinematic motion clips with precise control and audio sync

kling-3.0/4k/image-to-video

Animate stills into native 4K cinematic clips with start-end frame guidance and synchronized sound.

hunyuan/text-to-video

Turn text prompts into high quality videos with Tencent Hunyuan Video.

live-avatar

Turn still portraits into expressive, lifelike videos with control and precision.

wan-2-6/video-to-video

Transforms reference clips into 1080p short videos with precise motion and voice alignment.

Frequently Asked Questions

Is FLUX 3 Video available now on RunComfy?

FLUX 3 Video is coming soon. This preview lets you explore the FLUX 3 Video experience while Black Forest Labs finalizes the unified multimodal model and its API. Until the FLUX 3 backend is switched on, the live input fields and limits reflect the model currently powering this preview.

What is FLUX 3 Video used for?

FLUX 3 Video is designed to turn prompts and optional image, video, and audio references into short cinematic clips with native audio. It targets ad creative, film previsualization, and branded storytelling where reference consistency and synchronized sound matter.

What does FLUX 3 Video improve compared to earlier video approaches?

FLUX 3 Video is positioned to bring Black Forest Labs' work into motion and audio under one multimodal foundation, with five officially confirmed tasks — text-to-video, image-to-video, video-to-video, video + audio continuation, and keyframe-to-video — plus native audio, multilingual dialogue, and clips up to 20 seconds in a single generation. Exact gains depend on your content, so compare clips on the same prompt once the model goes live.

How long can a FLUX 3 Video clip be?

FLUX 3 Video generates clips up to 20 seconds in a single pass, roughly double the 10-second limit common to most video models, which leaves room for dialogue beats and slower dramatic shots. Clips can also be chained into multi-shot sequences that run for minutes while characters stay consistent. On the model currently powering this preview, duration runs from 4 to 15 seconds; the full FLUX 3 Video model is designed to reach 20 seconds.

Does FLUX 3 Video support native audio and lip-sync?

Yes. Native audio is a core goal for FLUX 3 Video, and on the model currently powering this preview you can enable audio generation to output synchronized speech, SFX, and music, which supports lip-synced, multilingual clips. You can disable it for silent video. Lip-sync quality depends on prompt clarity and the scene you describe.

Why do FLUX 3 Video motion and sound look realistic?

FLUX 3 was trained with video prediction as the bulk of its compute, so the model had to learn contact, motion, weight, and cause and effect, essentially how the physical world behaves. Native audio is learned on top of that world model, which is why sounds line up with the on-screen events that cause them and speech tracks lip movement. In practice, FLUX 3 Video is noted for expressive faces and physically plausible motion.

Can FLUX 3 Video keep characters or style consistent across a clip?

Reference guidance is central to FLUX 3 Video. In practice, reference images plus a clear prompt help anchor identity, wardrobe, and tone across frames, and FLUX 3 is designed to carry a subject from a source clip into a new scene. On the model currently powering this preview you can attach reference images, videos, and audio within the supported limits to reinforce consistency.

What input limits should I know before using FLUX 3 Video on RunComfy?

On the model currently powering this preview, prompts are recommended at Chinese ~≤500 characters or English ~≤1000 words. You can attach up to 9 reference images, 3 reference videos (2–15 seconds each), and 3 reference audio clips (2–15 seconds, under 15 MB). Only the prompt is required. Check the live parameter panel, since limits will change once FLUX 3 Video goes live.

What resolution, aspect ratio, and duration options does this FLUX 3 Video preview expose?

Resolution options are 480p, 720p (default), 1080p, and 4K. Aspect ratio can be adaptive (default, where the model picks the closest ratio) or fixed at 16:9, 9:16, 4:3, 3:4, 1:1, or 21:9. Duration is a whole number of seconds from 4 to 15 on the model currently powering this preview, with a default of 5; the full FLUX 3 Video model is designed to reach 20 seconds in a single generation.

Will FLUX 3 Video offer open weights or LoRA fine-tuning?

FLUX 3 AI is planned as one foundation with several editions. Alongside hosted FLUX 3 Video, Black Forest Labs has announced FLUX 3 Image, FLUX 3 Action (the FLUX-mimic robotics work), and FLUX 3 Dev, an open-weight multimodal backbone. Black Forest Labs has confirmed open weights are planned; once the FLUX 3 Dev backbone lands, expect a community FLUX 3 LoRA workflow, including FLUX 3 dev LoRA training for recurring characters, styles, and brand looks.

How do I move from testing FLUX 3 Video in the browser to production API integration?

Prototype in the RunComfy model UI, then call the same template through the RunComfy API using identical Input fields (prompt, images, videos, audios, aspect_ratio, duration, resolution, generate_audio, seed). Validate prompts and media limits in the UI first, then use your account API key and credits for automated jobs.

How much does it cost to generate on this FLUX 3 Video preview?

Generations consume credits based on the generated video duration on the model currently powering this preview, Seedance 2.0 Pro: $0.08 per second for 480p, $0.175 per second for 720p, $0.40 per second for 1080p, and $0.90 per second for 4K. See the pricing shown on this page, which will update when FLUX 3 Video launches.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Wan 2.6 Flash
  • Wan 2.6 Text to Video
  • Hailuo 2.3 Fast Standard
  • Happy Horse 1.1 reference to video
  • Seedance 2.5 Reference to Video
  • Seedance 2.0 Pro
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • seedream 4.0
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • Nano Banana 2 Edit
  • GPT Image 2 Image Edit
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of FLUX 3 Video

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...