Cinematic 4K image-to-video at $0.47 per second of output.
Gemini Omni Flash is Google's multimodal video generation family, built on Gemini's real-world knowledge and a stronger grasp of physics for more believable motion and interaction. The wider family spans four related tasks: text-to-video, image-to-video, video editing, and reference-to-video. This RunComfy page runs the text-to-video task, turning a written prompt into a short cinematic clip with synchronized audio.
Because it reasons about how the world actually behaves, the model keeps objects, lighting, and movement coherent across a shot instead of drifting frame to frame.
Inputs match the RunComfy OpenAPI Input schema for the text-to-video task.
| Parameter | Required | Type | Default | Range / Options | Description |
|---|---|---|---|---|---|
| prompt* | Yes (*) | string | — | descriptive text | The text prompt describing the video to generate |
| aspect_ratio | No | string | 16:9 | 16:9, 9:16 | Aspect ratio of the generated video |
| duration | No | integer | 8 | 3-10 (seconds) | Length of the generated video in whole seconds |
Billing is based on the length of the generated clip:
Cinematic 4K image-to-video at $0.47 per second of output.
Generate 768p, 2K video from image, video, and audio references
HappyHorse 1.0 Reference to Video fuses up to 9 reference images and a prompt into a coherent multi-character clip with stable identity.
Prompt-driven video editing at $0.126 per second of output.
Create dynamic, sound-synced motion clips from visuals for rich storytelling.
Convert visuals to cinematic videos quickly with Veo 3.1 Fast image-to-video for seamless creative control.
Gemini Omni Flash is Google's multimodal video model that turns a text prompt into a short cinematic clip with synchronized audio. On this RunComfy page it runs the text-to-video task, making it a fit for social ads, story beats, and quick previsualization where motion and sound both matter.
The Gemini Omni Flash family spans four related tasks: text-to-video, image-to-video, video editing, and reference-to-video. This RunComfy page exposes the text-to-video task, which generates a video from a written prompt alone, while the other tasks start from existing images or footage.
Yes. Gemini Omni Flash produces synchronized audio—such as speech, sound effects, and music—aligned to the generated action. You can steer the sound from the prompt, for example by asking for calm background music or specifying no dialogue.
Gemini Omni Flash is grounded in Gemini's real-world knowledge and has an improved understanding of physics, which helps it keep motion, lighting, and object interaction coherent across a shot. That reduces the frame-to-frame drift common in older text-to-video models.
The text-to-video task takes a required prompt, an aspect ratio of 16:9 or 9:16, and a duration between 3 and 10 seconds. Check the current RunComfy parameter panel for the exact defaults and any provider-side limits before you generate.
Be descriptive about subject, action, camera, mood, and lighting, and control pacing directly in the prompt (for example, a single continuous shot). Put exclusions in the prompt itself, such as "do not show text," since Gemini Omni Flash reads negative instructions from the prompt.
Yes. You can prototype Gemini Omni Flash in the RunComfy model UI and then call the same model via the RunComfy API with identical parameters. That lets you move from a browser test to automated generation without hosting or scaling the model yourself.
Generations with Gemini Omni Flash are billed at $0.13 per second of generated video, and they draw down your RunComfy usd or credit balance. New users typically start with a free trial amount; see the Generation section on this page for current details.
RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.





