logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

MiniMax H3 Max text to video: Prompt to Clip on Models and API | RunComfy

minimax/minimax-h3-max/text-to-video

MiniMax H3 Max text to video creates 5-15s clips from a prompt alone—often a 10s video in roughly a dozen seconds—with native audio at 480p or 768p.

Text prompt for video generation: scene, motion, camera, style, and optional audio.
The duration of the video in seconds.
The native generation resolution of the video.
The aspect ratio of the generated video.
How much effort to spend rewriting the prompt before generation. balanced returns quickly; quality spends longer on a richer prompt.
Random seed. A random seed is selected when omitted. Use -1 for random.
If set to true, the safety checker will be enabled.
Idle
The rate is $0.055 per second for 480p, and $0.088 per second for 768p.

Introduction To MiniMax H3 Max text to video

MiniMax's MiniMax H3 Max text to video turns a shot list into clips with native audio at 480p or 768p, often a 10-second video in roughly a dozen seconds.
Trading long waits and late-stage sound design for one prompt that covers motion, camera, and audio, MiniMax H3 Max text to video helps marketers, previz artists, and social teams iterate faster.
For developers, MiniMax H3 Max text to video on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: Scene Previz | Ad Storyboards | Social Hooks

MiniMax / MiniMax H3 Max Text To Video#


MiniMax H3 Max text to video is the text-only path of MiniMax's post-trained H3 Max line—and it is built for speed. A 10-second clip often comes back in roughly a dozen seconds, so you can draft, revise, and lock shots in a tight loop.


You describe the shot, camera, and sound; the model returns a short MP4 with matching audio at 480p or 768p. Use MiniMax H3 Max text to video when you want to invent framing from words rather than locking a still.


Highlights#


  • Fast turnaround: MiniMax H3 Max text to video often generates a 10-second video in roughly a dozen seconds, keeping iteration close to real time.
  • Text-only entry: No reference image required—MiniMax H3 Max text to video invents the first frame from your brief.
  • Prompt adherence: MiniMax H3 Max text to video keeps named beats, camera moves, and timing closer to the written order.
  • Native audio: Steer ambience, foley, or dialogue cues in the same MiniMax H3 Max text to video prompt as the picture.
  • Aspect control: MiniMax H3 Max text to video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 (default 16:9).
  • Prompt expansion: Use balanced for a fast rewrite or quality for a deeper expansion before generation.
  • Length & resolution: 5-15 seconds at 480p or 768p for drafts and delivery previews with MiniMax H3 Max text to video.

Parameters#


ParameterRequiredTypeDefaultRange / OptionsDescription
prompt *Yes (*)stringSample briefTextScene, motion, camera, style, and optional audio direction.
durationNointeger55-15Clip length in whole seconds.
resolutionNostring768p480p, 768pNative generation resolution.
aspect_ratioNostring16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Aspect ratio of the generated video.
prompt_expansion_mode *Yes (*)stringbalancedbalanced, qualityHow much effort to spend rewriting the prompt before generation.
seedNointegerrandomIntegerFixed seed for reproducible results; omit for a random seed.
enable_safety_checkerNobooleantruetrue, falseEnables the safety checker when true.

Pricing#


ResolutionPrice
480p$0.055 per second
768p$0.088 per second

Billing follows the generated clip duration. Check the Generation section on this page for the live credit estimate before you run MiniMax H3 Max text to video.


Related Models

kling-1-6/pro/text-to-video

Generate high quality videos from text prompts using Kling 1.6 Pro.

seedance-2.5/first-last-frame/720p

Bridge start and end stills into smooth cinematic video

kling-2-5/turbo/image-to-video

Render fluid, stylized scenes with fast, frame-consistent output

infinite-talk/image-to-video

Create photo-based, speech-aligned videos with natural motion

sora-2/image-to-video

Create lifelike scenes with synced audio and visual fidelity.

wan-2-5/text-to-video

Generate videos from text prompts with audio using Wan 2.5 Preview.

Frequently Asked Questions

What is MiniMax H3 Max text to video best at?

MiniMax H3 Max text to video turns a written shot list into a short clip with matching audio, without needing a reference image. It fits marketers, previz artists, and social teams who want motion, camera, and sound from one prompt.

How is MiniMax H3 Max text to video different from the image-to-video page?

MiniMax H3 Max text to video invents the first frame from text alone and exposes aspect-ratio controls. The image-to-video page locks an opening still and can optionally steer a last frame; pick the text path when framing should come from the brief.

What aspect ratios does MiniMax H3 Max text to video support?

MiniMax H3 Max text to video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, with 16:9 as the default. Choose the ratio that matches your delivery channel before generating.

What resolution and duration limits apply to MiniMax H3 Max text to video?

MiniMax H3 Max text to video generates 5 to 15 second clips at 480p or 768p. Use 480p for cheaper iteration and 768p for sharper previews; confirm the live options in the RunComfy parameter panel.

Does MiniMax H3 Max text to video include native audio?

Yes. MiniMax H3 Max text to video renders synchronized audio with the picture. Describe ambience, foley, or dialogue in the same prompt so sound and motion stay aligned in one pass.

What does prompt expansion do in MiniMax H3 Max text to video?

MiniMax H3 Max text to video offers balanced and quality expansion modes. Balanced returns a quick rewrite; quality spends longer enriching the prompt. Sparse briefs often benefit from quality; detailed shot lists usually work well with balanced.

Can developers call MiniMax H3 Max text to video via the RunComfy API?

Yes. Prototype MiniMax H3 Max text to video in the RunComfy Web UI, then reuse the same parameters through the RunComfy API for automation. Hosting and scaling stay on RunComfy's side.

How much does MiniMax H3 Max text to video cost on RunComfy?

MiniMax H3 Max text to video is billed per second: $0.055 per second at 480p and $0.088 per second at 768p. Generations consume USD or credits; check the Generation section on this page for the current estimate.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Wan 3.0 Prime Reference To Video
  • Wan 3.0 Prime
  • Wan 3.0 Prime Text To Video
  • Wan 2.6 Flash
  • MiniMax H3 Open Image to Video
  • MiniMax H3 Open
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • Qwen Image 3.0 Pro Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2 Image Edit
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of MiniMax H3 Max text to video

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...