Crea scene video realistiche da foto con movimento, sincronizzazione dei dialoghi e controllo creativo flessibile.
MiniMax H3 is MiniMax's multimodal model family for generating and editing images, video, and audio through natural-language instructions. This page provides MiniMax H3 Text-to-Video: turn one written prompt into a 4–15-second video at 24 FPS, at either 768p or 2K, with native stereo audio generated with the picture. It accepts text only; use the linked H3 workflows when you need image, video, or audio references.
| MiniMax H3 advantage | What it means for you |
|---|---|
| Selectable 768p or 2K with native stereo audio | Draft faster at 768p, or generate detailed visuals, dialogue, effects, ambience, and music in one pass at 2K—reducing separate upscaling, sound-design, and synchronization stages. |
| Language-directed Omni-Reference | Across MiniMax H3 workflows, assign images, video, and audio different roles—such as identity, product appearance, motion, camera style, voice, or music—so several sources can guide one coherent result. |
| V2V transfer and targeted editing | MiniMax H3 can carry over motion or camera language and revise selected visual or audio elements while preserving the rest, giving production teams a path beyond regenerating a shot from scratch. |
| Delivery-ready timing and framing | Generate 4–15 seconds at 24 FPS in six aspect ratios, so hooks, product reveals, and multi-beat scenes fit cinematic, web, feed, portrait, or vertical placements with less recutting and recropping. |
9:16 Shorts, Reels, and Stories or 1:1 feed assets with a clear opening hook and platform-ready framing.768p or 2k resolution.The first table lists the controls exposed by the MiniMax H3 Text-to-Video tool on this page.
| Parameter | Required | Type | Default | Range / Options | How to choose |
|---|---|---|---|---|---|
prompt* | Yes (*) | String | Example prompt | 1–4,000 characters | Give MiniMax H3 a structured brief covering the subject, one main action, camera, setting and lighting, visual style, audio, and intended ending. Clear structure matters more than using the full limit. |
aspect_ratio | No | String | 16:9 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Match MiniMax H3 output to the destination: 21:9 for ultra-wide cinematic shots, 16:9 for general video and ads, 4:3 for editorial or retro framing, 1:1 for feeds, 3:4 for portrait products, and 9:16 for vertical social video. |
resolution | No | String | 768p | 768p, 2k | Use 768p for cheaper drafts and faster iteration; use 2k when the shot needs higher detail for large placements or crops. |
duration | No | Integer | 5 | 4–15 seconds, in 1-second steps | Use 4–7-second MiniMax H3 clips for one action or fast iteration, 8–10 seconds for a simple change, and 11–15 seconds only when the prompt defines a clear beginning, development, and ending. |
The following table summarizes the wider MiniMax H3 model family. Inputs described for first/last-frame and Omni-Reference modes are available through separate workflows, not as controls on this Text-to-Video page.
| Core dimension | MiniMax H3 |
|---|---|
| Model | MiniMax-H3 |
| Output duration | MiniMax H3 outputs 4–15 seconds. |
| Output aspect ratio | First/last-frame mode: follows the original aspect ratio of the input image.<br>Text-to-Video mode: follows the user-selected 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 ratio.<br>Omni-Reference mode: uses one of those six ratios or Auto, which lets MiniMax H3 choose the output ratio. |
| Resolution | MiniMax H3 supports two tiers. 768p: for aspect ratios from 16:9 through 9:16, the short side is 768 pixels; for wider output, total resolution is about 1 MP—for example, 21:9 is 1536×672.<br>1440p / 2K: for aspect ratios from 16:9 through 9:16, the short side is 1440 pixels; for wider output, total resolution is about 3.7 MP—for example, 21:9 is 2976×1248. |
| Output frame rate | MiniMax H3 outputs at 24 FPS. |
| Output audio | Every MiniMax H3 result includes native stereo audio. |
| First/last-frame input | Images: 0, 1, or 2; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5). With no image input, MiniMax H3 runs in Text-to-Video mode—the tool provided on this page. |
| Omni-Reference input | Images: up to 9; width and height each 256–5760 pixels.<br>Videos: up to 3; each 2–15 seconds; combined video duration up to 15 seconds; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5).<br>Audio: up to 3 clips; each 2–15 seconds; combined audio duration up to 15 seconds. Audio must accompany an image or video and cannot be the only reference.<br>Mixed input: up to 12 files total. With no image, video, or audio input, the workflow becomes Text-to-Video. |
| Supported input formats | MiniMax H3 accepts video: H.264/AVC or H.265/HEVC; embedded audio: AAC or MP3.<br>Image: JPG, JPEG, PNG, WEBP, HEIC, or HEIF.<br>Audio: WAV or MP3. |
| Input size limits | Each video: 50 MB; each image: 30 MB; each audio file: 15 MB. There is no separate combined-media limit beyond the per-file limits, but the API request body is limited to 64 MB; URL-based media input is recommended. |
| Prompt limit | The MiniMax H3 product guide allows up to 7,000 characters. This Text-to-Video deployment currently accepts 1–4,000 characters, as shown in the page-control table above. |
MiniMax H3 Text-to-Video pricing depends on resolution and duration:
| Resolution | Price per second | 5s | 10s | 15s |
|---|---|---|---|---|
| 768p | $0.11 | $0.55 | $1.10 | $1.65 |
| 2K | $0.16 | $0.80 | $1.60 | $2.40 |
For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.
Use this reusable MiniMax H3 prompt structure:
[subject + defining details] + [one action over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [dialogue / effects / ambience / music] + [ending frame or constraint]
Weak prompt
> A luxury watch ad, cinematic and dynamic, with music.
Improved MiniMax H3 prompt
> A brushed-steel automatic watch rests on black volcanic stone. A narrow studio light sweeps across the sapphire crystal as the second hand moves and condensation beads on the case. Begin with an extreme macro, then make a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Audio: soft mechanical ticking and one low cinematic pulse; no dialogue.
The improved version gives MiniMax H3 an identifiable subject, timed action, separate camera direction, lighting, finish, sound sources, and a final frame that can be reviewed.
Use this MiniMax H3 comparison as a model-family guide; maximum resolution and duration may not be available together, and RunComfy workflows can expose different inputs, resolution tiers, and audio controls.
| Model | Resolution | Max duration | Audio | Standout |
|---|---|---|---|---|
| MiniMax H3 | 768p or 2K | 15s | Native stereo (voice, SFX, music) | Relates text, image, video, and audio references through language for V2V motion transfer and production editing; public weights are planned but have not yet been released. |
| Hailuo 02 | 1080p | ~10s | None | Strong prompt adherence and physics-focused motion suit gymnastics, dance, product movement, and other silent action shots that will receive audio later in production. |
| Kling 3.0 | Up to 4K | 15s | Native audio with lip-sync | Coordinates multi-shot camera changes with multilingual, speaker-assigned dialogue and lip-sync—useful for scripted ads, storyboards, and character-led scenes. |
| Seedance 2.0 | 1080p | 15s | Joint audio and video | Combines dense image, video, and audio references with joint audiovisual generation and precise lip-sync for identity-sensitive ads, branded stories, and reference-heavy edits. |
| Veo 3.1 | 4K | 8s | Native dialogue and effects | Pairs prompted dialogue and effects with first/last-frame, reference-image, and scene-extension controls—suited to cinematic transitions and assembled sequences. |
For Hailuo 02, the listed maxima are mode-specific: 1080p output is limited to 6 seconds, while 10-second output is available at 512p or 768p.
For MiniMax H3, Omni-Reference and V2V capabilities belong to sibling workflows; the Text-to-Video tool on this page remains prompt-only.
What sets MiniMax H3 apart is the combination: one general-purpose model family that relates text, image, video, and audio, generates native stereo sound, and offers both 768p ($0.11/s) and 2K ($0.16/s). Choose MiniMax H3 when you want to start from a text brief and keep a path to reference-guided creation or editing in sibling workflows. Based on publicly available information, run the same brief through each model before committing a pipeline.
If MiniMax H3 is not the right starting point for a project, compare these focused workflows on RunComfy:
Crea scene video realistiche da foto con movimento, sincronizzazione dei dialoghi e controllo creativo flessibile.
Anima l'immagine di un personaggio di riferimento con il movimento di un video di movimento utilizzando i controlli di suggerimento, risoluzione, inferenza, guida dell'immagine e guida della posa.
Strumento di conversione del movimento basato sull'intelligenza artificiale che consente la creazione di animazioni precise e stabili
Trasforma immagini statiche in clip video cinematografiche con transizioni fluide, realistiche e flessibilità creativa – powered by Seedance 1.5 Pro.
Crea un video con Wan 2.1 da un prompt di testo e imposta risoluzione, proporzioni, numero di fotogrammi, frame rate e altre opzioni.
Genera un video da un prompt obbligatorio con una durata di 6, 8 o 10 secondi, da 1080p a 2160p, 25 o 50 FPS, proporzioni 16:9 fisse e audio opzionale.
SÌ. MiniMax H3 Text-to-Video è disponibile nel browser e tramite l'API RunComfy utilizzando i crediti del tuo account.
MiniMax H3 è la famiglia di modelli di generazione e modifica multimodali per uso generale di MiniMax. La sua progettazione più ampia delle attività copre testo, immagini, video e audio, mentre i singoli flussi di lavoro espongono input diversi. Questa pagina fornisce il flusso di lavoro di sola richiesta da testo a video.
MiniMax posiziona H3 per pubblicità, branding, e-commerce, film, design di titoli, poster animati, narrazione in forma breve, concetti di prodotto e interfaccia utente, giochi, personaggi virtuali e animazioni stilizzate. L'attuale flusso di lavoro da testo a video è ideale per concetti brevi che possono essere indirizzati da un suggerimento scritto.
In questa pagina, MiniMax H3 supporta la risoluzione "768p" e "2k". Scegli qualsiasi durata di secondo intero compresa tra 4 e 15 secondi; il valore predefinito è 5 secondi. Usa 768p per bozze più economiche e 2K quando hai bisogno di dettagli più elevati.
SÌ. Questo flusso di lavoro da testo a video genera audio stereo nativo con il video. Descrivi il dialogo, l'atmosfera, gli effetti sonori o la musica nel prompt; non è disponibile un interruttore audio separato. I risultati possono variare, quindi rivedi la sincronizzazione labiale e il timing audio.
No. Questa pagina è il flusso di lavoro di sola richiesta testo-video MiniMax H3 e la sua API accetta solo "prompt", "aspect_ratio", "risoluzione" e "durata". Per il supporto sorgente, utilizzare un flusso di lavoro H3 Image-to-Video o Reference-to-Video separato.
Nei flussi di lavoro H3 abilitati per i riferimenti, il linguaggio naturale può assegnare ruoli diversi a testo, immagini, video e audio. Ad esempio, una fonte potrebbe definire il movimento della telecamera, un'altra il personaggio e un'altra la voce. Questa è una funzionalità della famiglia di modelli, non una funzionalità di caricamento di contenuti multimediali nella pagina corrente.
No. RunComfy fornisce MiniMax H3 tramite il browser e un'API HTTP, quindi non è necessario ospitare o ridimensionare il modello da soli. Il post di lancio di MiniMax del 31 luglio descriveva un piano condizionato per rilasciare i pesi del modello; verificare la disponibilità attuale e i termini di licenza prima di pianificare il self-hosting.
MiniMax H3 Text-to-Video costa $ 0,11 per secondo generato a 768p e $ 0,16 per secondo generato a 2K. A 768p, un video di 5 secondi costa $ 0,55, un video di 10 secondi costa $ 1,10 e un video di 15 secondi costa $ 1,65. A 2K, quelle lunghezze costano $ 0,80, $ 1,60 e $ 2,40. La generazione batch moltiplica il costo basato sulla durata per il numero di output.
MiniMax descrive Hailuo 02 come incentrato su architettura, dati e scala, mentre MiniMax H3 si concentra sulla generalizzazione di compiti e modalità. H3 aggiunge inoltre audio stereo nativo e output selezionabile 768p o 2K al flusso di lavoro da testo a video.
Testa il messaggio, le proporzioni, la risoluzione e la durata in RunComfy. Quindi chiama lo stesso modello Text-to-Video MiniMax H3 tramite l'API con "prompt" (richiesto), "aspect_ratio", "risoluzione" (768p o "2k") e "durata" (4–15). Questo flusso di lavoro non ha campi di caricamento multimediale.
RunComfy è la piattaforma principale ComfyUI che offre ComfyUI online ambiente e servizi, insieme a workflow di ComfyUI con visuali mozzafiato. RunComfy offre anche AI Models, consentire agli artisti di sfruttare gli ultimi strumenti di AI per creare arte incredibile.














