logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

Gemini Omni Flash: Text-to-Video with Synced Audio on Models and API | RunComfy

google/gemini-omni-flash/text-to-video

Gemini Omni Flash generates short cinematic videos with native synchronized audio from a text prompt, supporting 16:9 or 9:16 framing and 3 to 10 second durations.

The text prompt describing the video you want to generate. You can control pacing and audio directly in the prompt (e.g. "in a single continuous shot", "include calm background music", "no dialogue") and add negative instructions such as "Do not show text".
The aspect ratio of the generated video.
The duration of the generated video, in seconds.
Idle
The rate is $0.13 per second of generated video.

Introduction To Gemini Omni Flash

Google's Gemini Omni Flash turns a single text prompt into a cinematic 3-10 second clip with synchronized audio, grounded in Gemini's real-world knowledge for coherent, physics-aware motion. Trading manual editing, separate sound design, and multi-tool pipelines for one prompt-driven step, it serves marketers, filmmakers, game teams, and social creators. For developers, Gemini Omni Flash on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: Short-Form Social Ads | Cinematic Story Beats | Rapid Previsualization

Google / Gemini Omni Flash#


Gemini Omni Flash is Google's multimodal video generation family, built on Gemini's real-world knowledge and a stronger grasp of physics for more believable motion and interaction. The wider family spans four related tasks: text-to-video, image-to-video, video editing, and reference-to-video. This RunComfy page runs the text-to-video task, turning a written prompt into a short cinematic clip with synchronized audio.


Because it reasons about how the world actually behaves, the model keeps objects, lighting, and movement coherent across a shot instead of drifting frame to frame.


Highlights#


  • Native synchronized audio: Generate speech, sound effects, and music aligned to the action, so a clip is ready to share without a separate audio pass.
  • Physics-aware motion: Improved physical understanding produces steadier, more natural movement and object interaction.
  • Grounded in real knowledge: Outputs draw on Gemini's real-world knowledge, helping scenes feel plausible and on-topic.
  • Prompt-level control: Direct pacing and sound straight from the prompt (for example, a single continuous shot, calm background music, or no dialogue).
  • Four-task family: The Gemini Omni Flash lineup covers text-to-video, image-to-video, video editing, and reference-to-video; this page exposes the text-to-video task alone.
  • Flexible framing: Supports 16:9 landscape and 9:16 portrait, with 3 to 10 second durations.

Parameters#


Inputs match the RunComfy OpenAPI Input schema for the text-to-video task.


ParameterRequiredTypeDefaultRange / OptionsDescription
prompt*Yes (*)string—descriptive textThe text prompt describing the video to generate
aspect_ratioNostring16:916:9, 9:16Aspect ratio of the generated video
durationNointeger83-10 (seconds)Length of the generated video in whole seconds

Pricing#


Billing is based on the length of the generated clip:


  • Video: $0.13 per second of generated video.

How to Use#


  1. Write a descriptive prompt — Describe subject, action, setting, mood, lighting, and camera. Detailed prompts give the model the most to work with.
  2. Set the pacing in words — Ask for a single continuous shot or a specific rhythm directly in the prompt.
  3. Direct the audio — Request dialogue, ambient sound, or background music, or say no dialogue when you want it quiet.
  4. Add negative instructions — Put exclusions in the prompt itself, such as do not show text.
  5. Choose an aspect ratio — Pick 16:9 for landscape or 9:16 for vertical, mobile-first delivery.
  6. Set the duration — Choose any whole number of seconds from 3 to 10.
  7. Generate and refine — Review the clip, then adjust wording, ratio, or duration and run again.

Prompt & Reference Tips#


  • Lead with the camera and motion you want (for example, slow push-in, handheld, or locked-off).
  • Name the mood and lighting so the scene reads the way you intend.
  • Keep one clear action per clip; short durations reward focused prompts.
  • Spell out the soundscape when audio matters, since the model can generate synchronized audio.
  • Use plain negative phrasing (for example, no on-screen text) rather than vague hints.
  • If a result feels busy, simplify the prompt and remove competing style cues.

How Gemini Omni Flash compares to other models#


  • Versus earlier text-to-video models: It emphasizes physics-aware motion and native synchronized audio in one step. Based on publicly available information, compare on your own prompts.
  • Versus silent video generators: It bundles audio generation, so you skip a separate scoring or sound-design stage.
  • Within its own family: Text-to-video is one of four tasks in this family; image-to-video, video editing, and reference-to-video target workflows that start from existing media instead of text alone.

More Models to Try#


  • Veo 3 Fast — Fast text-to-video with audio for quick drafts.
  • Seedance 2.5 — Cinematic multimodal video with reference support.
  • Kling Video — Motion-focused video generation.
  • Wan 2.5 — General-purpose text-to-video alternative.

Official Resources#


  • Google DeepMind: https://deepmind.google/

Related Models

seedance-2.5/image-to-video

Seedance 2.5 Coming Soon: Animate a still image into cinematic AI video

hailuo-2-3/standard/text-to-video

Create expressive AI videos from prompts with smooth motion and vivid detail.

kling-video-o3/pro/text-to-video

Cinematic Pro-tier text-to-video at $0.112 per second of output.

seedance-1.0/image-to-video

Create fluid, expressive animations with multi-shot storytelling features.

seedance-1.0/pro-fast/text-to-video

High-speed text-to-motion generator for cinematic storytelling use.

seedance-1.0/text-to-video

Generate cinematic videos from text prompts with Seedance 1.0.

Frequently Asked Questions

What is Gemini Omni Flash used for?

Gemini Omni Flash is Google's multimodal video model that turns a text prompt into a short cinematic clip with synchronized audio. On this RunComfy page it runs the text-to-video task, making it a fit for social ads, story beats, and quick previsualization where motion and sound both matter.

Which tasks does the Gemini Omni Flash family cover?

The Gemini Omni Flash family spans four related tasks: text-to-video, image-to-video, video editing, and reference-to-video. This RunComfy page exposes the text-to-video task, which generates a video from a written prompt alone, while the other tasks start from existing images or footage.

Does Gemini Omni Flash generate audio with the video?

Yes. Gemini Omni Flash produces synchronized audio—such as speech, sound effects, and music—aligned to the generated action. You can steer the sound from the prompt, for example by asking for calm background music or specifying no dialogue.

What makes Gemini Omni Flash's motion look more natural?

Gemini Omni Flash is grounded in Gemini's real-world knowledge and has an improved understanding of physics, which helps it keep motion, lighting, and object interaction coherent across a shot. That reduces the frame-to-frame drift common in older text-to-video models.

What input limits should I know before using Gemini Omni Flash?

The text-to-video task takes a required prompt, an aspect ratio of 16:9 or 9:16, and a duration between 3 and 10 seconds. Check the current RunComfy parameter panel for the exact defaults and any provider-side limits before you generate.

How should I write prompts for Gemini Omni Flash?

Be descriptive about subject, action, camera, mood, and lighting, and control pacing directly in the prompt (for example, a single continuous shot). Put exclusions in the prompt itself, such as "do not show text," since Gemini Omni Flash reads negative instructions from the prompt.

Can developers use Gemini Omni Flash through the RunComfy API?

Yes. You can prototype Gemini Omni Flash in the RunComfy model UI and then call the same model via the RunComfy API with identical parameters. That lets you move from a browser test to automated generation without hosting or scaling the model yourself.

How much does it cost to generate with Gemini Omni Flash on RunComfy?

Generations with Gemini Omni Flash are billed at $0.13 per second of generated video, and they draw down your RunComfy usd or credit balance. New users typically start with a free trial amount; see the Generation section on this page for current details.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Gemini Omni Flash Video Edit
  • Gemini Omni Flash Image to Video
  • Gemini Omni Flash Reference to Video
  • Wan 2.6 Flash
  • Seedance 2.0 Pro
  • Wan 2.7 Reference to Video
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • seedream 4.0
  • Nano Banana 2 Edit
  • Nano Banana Pro
  • GPT Image 2 Image Edit
  • GPT Image 2
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of Gemini Omni Flash

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...