logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

MiniMax H3 Max Reference to video: Multimodal Video on Models and API | RunComfy

minimax/minimax-h3-max/reference-to-video

MiniMax H3 Max Reference to video turns image, video, and audio references plus a prompt into 5-15s 480p/768p clips with native audio.

Describe the scene, motion, camera, and audio. Refer to assets as Image 1, Image 2, Video 1, Audio 1, and say what each reference supplies.
Image 1
Image 2
Subject or style reference image URLs, cited in the prompt as Image 1, Image 2, and so on. Combined with videos and audios, at most 12 reference files.
Motion reference clips (about 2-15s each; combined duration at most 15s), cited as Video 1, Video 2. Combined with images and audios, at most 12 reference files.
Optional voice or ambience clips (about 2-15s each; combined duration at most 15s). Cannot be the only reference; provide at least one image or video with them.
Output aspect ratio. adaptive follows the reference framing when possible.
Native generation resolution. 480p is faster/cheaper; 768p is the sharper default.
Length of the generated video in seconds (5-15).
How much effort to spend rewriting the prompt before generation. balanced returns quickly; quality spends longer on a richer prompt.
Fixed seed for reproducible results. Use -1 for a random seed.
If set to true, the safety checker will be enabled.
Idle
Output video is billed at $0.09 per second, plus $0.03 per reference image. Example: a 5s clip with 2 reference images costs $0.51.

Introduction To MiniMax H3 Max Reference to video

MiniMax's MiniMax H3 Max Reference to video builds a 480p or 768p clip from a written brief plus the image, video, and audio references you attach, rendering picture and native audio in one pass.
Trading lookalike casting and manual motion matching for references the model reads directly, MiniMax H3 Max Reference to video helps brand teams, agencies, and previz artists hold a character, product, or camera language steady across shots.
For developers, MiniMax H3 Max Reference to video on RunComfy can be used both in the browser and via an HTTP API, so you don't need to host or scale the model yourself.
Ideal for: Character Consistency | Product Reveals | Style Matching

MiniMax / MiniMax H3 Max Reference To Video#


Parameters#


ParameterRequiredTypeDefaultRange / OptionsDescription
prompt *Yes (*)stringSample briefTextScene, motion, camera, and audio. Cite Image 1 / Video 1 / Audio 1 and say what each reference supplies.
reference_images *Yes (*)array (image URL)Sample stillUp to 12 combined refsSubject or style images cited as Image 1, Image 2, …
reference_videosNoarray (video URL)NoneUp to 12 combined refs; ~2-15s each, ≤15s combinedMotion clips cited as Video 1, Video 2, …
reference_audiosNoarray (audio URL)NoneUp to 12 combined refs; ~2-15s each, ≤15s combinedVoice or ambience; cannot be the only reference.
aspect_ratioNostringadaptiveadaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16Output framing.
resolutionNostring768p480p, 768pNative generation resolution.
durationNointeger55-15Clip length in whole seconds.
prompt_expansion_mode *Yes (*)stringbalancedbalanced, qualityPrompt rewrite effort before generation.
seedNointeger-1 (random)IntegerFixed seed for reproducibility.
enable_safety_checkerNobooleantruetrue, falseEnables the safety checker when true.

Pricing#


ItemPrice
Output video$0.09 per second
Reference image$0.03 per image

Resolution does not change the per-second rate. Example: a 5s clip with 2 reference images costs $0.51. Check the Generation section on this page for the live credit estimate.

Related Models

kling/lipsync/text-to-video

Create lifelike speech-synced visuals from scripts or clips with Kling Lipsync for precise facial animation and realistic results.

kling-video-o1/image-to-video

Transform static visuals into cinematic motion with Kling O1's precise scene control and lifelike generation.

wan-2-2/image-to-video

Refined AI visuals, real-time control, and pro FX for creators

wan-2-1/fusionx/image-to-video

Cinema-grade AI videos with precise dual-prompt control

kling-2-1/standard/image-to-video

Animate a single image into a smooth video with Kling 2.1 Standard.

minimax-h3/text-to-video

MiniMax H3: 768p/2K text-to-video with native stereo audio

Frequently Asked Questions

What is MiniMax H3 Max Reference to video used for?

MiniMax H3 Max Reference to video generates short clips from a prompt plus image, video, and/or audio references. It is a strong fit when you need character, product, or style consistency across shots while picture and native audio are produced together.

How does MiniMax H3 Max Reference to video differ from plain image-to-video?

Image-to-video usually animates one opening frame. MiniMax H3 Max Reference to video conditions on multiple multimodal references—images for identity or look, optional videos for motion, optional audio for voice or ambience—so you can steer several cues in one brief.

What input limits should I know before using MiniMax H3 Max Reference to video?

Clips are 5–15 seconds at 480p or 768p. Reference images, videos, and audios together may total at most 12 files; video and audio refs are typically about 2–15 seconds each with combined duration up to 15 seconds. Audio cannot be the only reference—include at least one image or video. Check the parameter panel for live limits.

How should I write prompts for MiniMax H3 Max Reference to video?

Name each asset in order (Image 1, Video 1, Audio 1) and say what it supplies—identity, wardrobe, motion, or voice. Add clear camera and audio lines so MiniMax H3 Max Reference to video keeps lip-sync, ambience, and framing aligned with the references.

Does MiniMax H3 Max Reference to video generate audio with the picture?

Yes. MiniMax H3 Max Reference to video renders native audio in the same pass as the video. You can also attach reference audio when you need a specific voice or ambience, as long as an image or video reference is present too.

What resolutions and aspect ratios does MiniMax H3 Max Reference to video support?

MiniMax H3 Max Reference to video supports 480p and 768p, with aspect ratios including adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Adaptive follows the reference framing when possible.

Can developers use MiniMax H3 Max Reference to video through the RunComfy API?

Yes. Prototype MiniMax H3 Max Reference to video in the RunComfy Web UI, then call the same model and parameters over the RunComfy HTTP API for automation. New accounts typically receive a free trial USD balance to start testing.

How much does it cost to generate with MiniMax H3 Max Reference to video on RunComfy?

On RunComfy, MiniMax H3 Max Reference to video is billed at $0.09 per second of output video plus $0.03 per reference image. For example, a 5-second clip with 2 reference images costs $0.51. Generations consume USD/credits; check the Generation section on the model page for the live estimate.

Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Wan 3.0 Prime Reference To Video
  • Wan 3.0 Prime
  • Wan 3.0 Prime Text To Video
  • Wan 2.6 Flash
  • MiniMax H3 Open Image to Video
  • MiniMax H3 Open
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • Qwen Image 3.0 Pro Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2 Image Edit
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Examples Of MiniMax H3 Max Reference to video

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...