logo
RunComfy
  • ComfyUI
  • EntraîneurNouveau
  • Modèles
  • API
  • Tarification
discord logo
MODÈLES
Explorer
Tous les modèles
BIBLIOTHÈQUE
Générations
APIS DE MODÈLES
Documentation API
Clés API
COMPTE
Utilisation

MiniMax H3 : texte-vidéo 768p et 2K avec audio stéréo | CourirConfort | Models and API | RunComfy

minimax/minimax-h3/text-to-video

MiniMax H3 Text-to-Video transforme une invite écrite en une vidéo 768p ou 2K de 4 à 15 secondes avec un son stéréo natif. Utilisez un flux de travail H3 distinct pour les références image, vidéo ou audio.

Le rapport hauteur/largeur de la vidéo générée.
La résolution de la vidéo générée. Choisissez 768p pour des brouillons plus rapides/moins chers ou 2k pour plus de détails.
La durée de la vidéo générée en secondes (4 à 15).
Idle
The rate is $0.11 per second for 768p, and $0.16 per second for 2k.

Introduction à la création vidéo MiniMax H3

MiniMax H3 est la famille de modèles multimodaux à usage général de MiniMax. Cette page Texte en vidéo transforme une invite écrite en une vidéo de 4 à 15 secondes à 24 FPS avec un son stéréo généré à côté de l'image. Choisissez une résolution 768p ou 2K. Dans la famille plus large, les flux de travail H3 distincts utilisent les images, la vidéo et l'audio comme références pour l'identité, le mouvement, la caméra, la voix, le son ou l'édition ; ces entrées multimédias ne sont pas disponibles sur cette page.

Why Choose MiniMax H3#


MiniMax H3 is MiniMax's multimodal model family for generating and editing images, video, and audio through natural-language instructions. This page provides MiniMax H3 Text-to-Video: turn one written prompt into a 4–15-second video at 24 FPS, at either 768p or 2K, with native stereo audio generated with the picture. It accepts text only; use the linked H3 workflows when you need image, video, or audio references.


MiniMax H3 advantageWhat it means for you
Selectable 768p or 2K with native stereo audioDraft faster at 768p, or generate detailed visuals, dialogue, effects, ambience, and music in one pass at 2K—reducing separate upscaling, sound-design, and synchronization stages.
Language-directed Omni-ReferenceAcross MiniMax H3 workflows, assign images, video, and audio different roles—such as identity, product appearance, motion, camera style, voice, or music—so several sources can guide one coherent result.
V2V transfer and targeted editingMiniMax H3 can carry over motion or camera language and revise selected visual or audio elements while preserving the rest, giving production teams a path beyond regenerating a shot from scratch.
Delivery-ready timing and framingGenerate 4–15 seconds at 24 FPS in six aspect ratios, so hooks, product reveals, and multi-beat scenes fit cinematic, web, feed, portrait, or vertical placements with less recutting and recropping.

Best Use Cases#


  • Advertising and product launches: Use MiniMax H3 for hero shots, product reveals, campaign concepts, and short commercials where camera direction and sound should be designed together.
  • Social media content: Use MiniMax H3 to create 9:16 Shorts, Reels, and Stories or 1:1 feed assets with a clear opening hook and platform-ready framing.
  • Film and commercial previsualization: Use MiniMax H3 to test a camera move, lighting plan, action beat, scene transition, or temporary soundtrack before committing to production.
  • Brand and stylized storytelling: Use MiniMax H3 to explore title sequences, animated posters, character moments, and graphic looks when a written idea needs a polished audiovisual treatment.

How It Works#


  1. Write the shot: Tell MiniMax H3 what appears, what changes over time, how the camera moves, how the scene should look, and what should be heard.
  2. Choose the delivery format: Select an aspect ratio, a 4–15-second duration, and either 768p or 2k resolution.
  3. Generate and refine: The model creates video and stereo audio together. Review motion, text, faces, lip sync, and sound timing, then change one instruction at a time.

Parameters#


The first table lists the controls exposed by the MiniMax H3 Text-to-Video tool on this page.


ParameterRequiredTypeDefaultRange / OptionsHow to choose
prompt*Yes (*)StringExample prompt1–4,000 charactersGive MiniMax H3 a structured brief covering the subject, one main action, camera, setting and lighting, visual style, audio, and intended ending. Clear structure matters more than using the full limit.
aspect_ratioNoString16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Match MiniMax H3 output to the destination: 21:9 for ultra-wide cinematic shots, 16:9 for general video and ads, 4:3 for editorial or retro framing, 1:1 for feeds, 3:4 for portrait products, and 9:16 for vertical social video.
resolutionNoString768p768p, 2kUse 768p for cheaper drafts and faster iteration; use 2k when the shot needs higher detail for large placements or crops.
durationNoInteger54–15 seconds, in 1-second stepsUse 4–7-second MiniMax H3 clips for one action or fast iteration, 8–10 seconds for a simple change, and 11–15 seconds only when the prompt defines a clear beginning, development, and ending.

  • Required field.

The following table summarizes the wider MiniMax H3 model family. Inputs described for first/last-frame and Omni-Reference modes are available through separate workflows, not as controls on this Text-to-Video page.


Core dimensionMiniMax H3
ModelMiniMax-H3
Output durationMiniMax H3 outputs 4–15 seconds.
Output aspect ratioFirst/last-frame mode: follows the original aspect ratio of the input image.<br>Text-to-Video mode: follows the user-selected 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 ratio.<br>Omni-Reference mode: uses one of those six ratios or Auto, which lets MiniMax H3 choose the output ratio.
ResolutionMiniMax H3 supports two tiers. 768p: for aspect ratios from 16:9 through 9:16, the short side is 768 pixels; for wider output, total resolution is about 1 MP—for example, 21:9 is 1536×672.<br>1440p / 2K: for aspect ratios from 16:9 through 9:16, the short side is 1440 pixels; for wider output, total resolution is about 3.7 MP—for example, 21:9 is 2976×1248.
Output frame rateMiniMax H3 outputs at 24 FPS.
Output audioEvery MiniMax H3 result includes native stereo audio.
First/last-frame inputImages: 0, 1, or 2; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5). With no image input, MiniMax H3 runs in Text-to-Video mode—the tool provided on this page.
Omni-Reference inputImages: up to 9; width and height each 256–5760 pixels.<br>Videos: up to 3; each 2–15 seconds; combined video duration up to 15 seconds; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5).<br>Audio: up to 3 clips; each 2–15 seconds; combined audio duration up to 15 seconds. Audio must accompany an image or video and cannot be the only reference.<br>Mixed input: up to 12 files total. With no image, video, or audio input, the workflow becomes Text-to-Video.
Supported input formatsMiniMax H3 accepts video: H.264/AVC or H.265/HEVC; embedded audio: AAC or MP3.<br>Image: JPG, JPEG, PNG, WEBP, HEIC, or HEIF.<br>Audio: WAV or MP3.
Input size limitsEach video: 50 MB; each image: 30 MB; each audio file: 15 MB. There is no separate combined-media limit beyond the per-file limits, but the API request body is limited to 64 MB; URL-based media input is recommended.
Prompt limitThe MiniMax H3 product guide allows up to 7,000 characters. This Text-to-Video deployment currently accepts 1–4,000 characters, as shown in the page-control table above.

Pricing#


MiniMax H3 Text-to-Video pricing depends on resolution and duration:


ResolutionPrice per second5s10s15s
768p$0.11$0.55$1.10$1.65
2K$0.16$0.80$1.60$2.40

For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.


Prompting Tips & Examples#


Use this reusable MiniMax H3 prompt structure:


[subject + defining details] + [one action over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [dialogue / effects / ambience / music] + [ending frame or constraint]


  • Separate motion: Describe subject movement and camera movement independently.
  • Show sequence: Use “begins with,” “then,” and “ends on” when the clip has more than one beat.
  • Name sound sources: Tell MiniMax H3 who speaks, what makes each sound, the ambience, and whether music should lead or stay subtle.
  • Limit competing ideas: One main subject, one primary action, and one coherent style are easier for MiniMax H3 to follow.

Weak prompt


> A luxury watch ad, cinematic and dynamic, with music.


Improved MiniMax H3 prompt


> A brushed-steel automatic watch rests on black volcanic stone. A narrow studio light sweeps across the sapphire crystal as the second hand moves and condensation beads on the case. Begin with an extreme macro, then make a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Audio: soft mechanical ticking and one low cinematic pulse; no dialogue.


The improved version gives MiniMax H3 an identifiable subject, timed action, separate camera direction, lighting, finish, sound sources, and a final frame that can be reviewed.


How MiniMax H3 Compares#


Use this MiniMax H3 comparison as a model-family guide; maximum resolution and duration may not be available together, and RunComfy workflows can expose different inputs, resolution tiers, and audio controls.


ModelResolutionMax durationAudioStandout
MiniMax H3768p or 2K15sNative stereo (voice, SFX, music)Relates text, image, video, and audio references through language for V2V motion transfer and production editing; public weights are planned but have not yet been released.
Hailuo 021080p~10sNoneStrong prompt adherence and physics-focused motion suit gymnastics, dance, product movement, and other silent action shots that will receive audio later in production.
Kling 3.0Up to 4K15sNative audio with lip-syncCoordinates multi-shot camera changes with multilingual, speaker-assigned dialogue and lip-sync—useful for scripted ads, storyboards, and character-led scenes.
Seedance 2.01080p15sJoint audio and videoCombines dense image, video, and audio references with joint audiovisual generation and precise lip-sync for identity-sensitive ads, branded stories, and reference-heavy edits.
Veo 3.14K8sNative dialogue and effectsPairs prompted dialogue and effects with first/last-frame, reference-image, and scene-extension controls—suited to cinematic transitions and assembled sequences.

For Hailuo 02, the listed maxima are mode-specific: 1080p output is limited to 6 seconds, while 10-second output is available at 512p or 768p.


For MiniMax H3, Omni-Reference and V2V capabilities belong to sibling workflows; the Text-to-Video tool on this page remains prompt-only.


What sets MiniMax H3 apart is the combination: one general-purpose model family that relates text, image, video, and audio, generates native stereo sound, and offers both 768p ($0.11/s) and 2K ($0.16/s). Choose MiniMax H3 when you want to start from a text brief and keep a path to reference-guided creation or editing in sibling workflows. Based on publicly available information, run the same brief through each model before committing a pipeline.


More Models to Try#


If MiniMax H3 is not the right starting point for a project, compare these focused workflows on RunComfy:


  • MiniMax H3 Image-to-Video: Start from an image and optionally guide the final frame when composition, identity, or product appearance must stay anchored.
  • MiniMax H3 Reference-to-Video: Combine images with optional video and audio when motion, camera style, voice, music, or other source details must carry over.
  • Hailuo 02 Image-to-Video: Try an earlier MiniMax generation when you want a focused image-animation workflow.
  • Hailuo 2.3 Pro: Animate a still image in 1080p with stable, physics-aware motion, nuanced facial expression, and natural lighting—well suited to polished product, portrait, and character shots.
  • Kling 3.0: Animate a start image with an optional end frame, element references, timed shot prompts, and synchronized audio for controlled brand and character sequences.
  • Seedance 2.0 Pro: Mix up to nine images, three videos, and three audio clips to keep identity, camera language, and synchronized speech, effects, and music aligned across 4–15-second clips.

Official Resources#


  • MiniMax H3 launch post
  • MiniMax official website

Modèles associés

kling-3.0/pro/image-to-video

Image-to-video premium offrant la meilleure fidélité visuelle et les mouvements les plus réalistes de la gamme Kling V3.0.

runway-aleph/video-to-video

Montages vidéo cinématiques avec contrôle du style et objets dynamiques

ai-avatar/v2/standard

Transformez un portrait et une piste audio en une vidéo d'avatar parlant, avec un prompt facultatif pour guider les mouvements ou la manière de parler.

happyhorse-1.0/reference-to-video

HappyHorse 1.0 Reference to Video fusionne jusqu'à 9 images de référence et une invite dans un clip cohérent à plusieurs caractères avec une identité stable. Use HappyHorse 1.0 Reference to Video on RunComfy.

one-to-all-animation/1.3b

Animez une image de personnage de référence avec le mouvement d'une vidéo pilote à l'aide d'invites positives et négatives ainsi que de paramètres de génération réglables.

hunyuan/image-to-video

Transitions fluides, mouvements cohérents et rendu vidéo de qualité.

Questions Fréquemment Posées

Le MiniMax H3 est-il disponible maintenant sur RunComfy ?

Oui. MiniMax H3 Text-to-Video est disponible dans le navigateur et via l'API RunComfy en utilisant les crédits de votre compte.

Qu’est-ce que le MiniMax H3 ?

MiniMax H3 est la famille de modèles de génération et d'édition multimodales à usage général de MiniMax. Sa conception de tâches plus large couvre le texte, l'image, la vidéo et l'audio, tandis que les flux de travail individuels exposent différentes entrées. Cette page fournit le flux de travail Texte-Vidéo avec invite uniquement.

À quoi peut servir MiniMax H3 ?

MiniMax positionne H3 pour la publicité, l'image de marque, le commerce électronique, le cinéma, la conception de titres, les affiches animées, la narration courte, les concepts de produits et d'interface utilisateur, les jeux, les personnages virtuels et l'animation stylisée. Le flux de travail texte-vidéo actuel est idéal pour les concepts courts pouvant être dirigés à partir d’une invite écrite.

Quelle résolution et quelle longueur de clip le MiniMax H3 Text-to-Video prend-il en charge ?

Sur cette page, MiniMax H3 prend en charge les résolutions « 768p » et « 2k ». Choisissez n’importe quelle durée d’une seconde entière comprise entre 4 et 15 secondes ; la valeur par défaut est de 5 secondes. Utilisez 768p pour des brouillons moins chers et 2K lorsque vous avez besoin de plus de détails.

Le MiniMax H3 génère-t-il de l'audio ?

Oui. Ce flux de travail Text-to-Video génère un son stéréo natif avec la vidéo. Décrivez le dialogue, l'ambiance, les effets sonores ou la musique dans l'invite ; il n’y a pas de bascule audio séparée. Les résultats peuvent varier, alors vérifiez la synchronisation labiale et le timing audio.

Puis-je utiliser des références d’images, de vidéos ou d’audio sur cette page MiniMax H3 ?

Non. Cette page est le flux de travail texte-vidéo MiniMax H3 à invite uniquement, et son API n'accepte que « invite », « aspect_ratio », « résolution » et « durée ». Pour le média source, utilisez un flux de travail image-vers-vidéo ou référence-vidéo H3 distinct.

Que signifie le contexte multimodal pour la famille de modèles MiniMax H3 ?

Dans les flux de travail H3 prenant en charge les références, le langage naturel peut attribuer différents rôles au texte, aux images, à la vidéo et à l'audio. Par exemple, une source peut définir le mouvement de la caméra, une autre le personnage et une autre la voix. Il s'agit d'une fonctionnalité de famille de modèles et non d'une fonctionnalité de téléchargement multimédia sur la page actuelle.

Dois-je auto-héberger le MiniMax H3 ?

Non. RunComfy fournit MiniMax H3 via le navigateur et une API HTTP, vous n'avez donc pas besoin d'héberger ou de mettre à l'échelle le modèle vous-même. Le message de lancement de MiniMax du 31 juillet décrivait un plan conditionnel pour publier les poids des modèles ; vérifiez la disponibilité actuelle et les conditions de licence avant de planifier l’auto-hébergement.

Combien coûte la synthèse texte-vidéo MiniMax H3 ?

La conversion texte-vidéo MiniMax H3 coûte 0,11 $ par seconde générée à 768p et 0,16 $ par seconde générée à 2K. À 768p, une vidéo de 5 secondes coûte 0,55 $, une vidéo de 10 secondes coûte 1,10 $ et une vidéo de 15 secondes coûte 1,65 $. À 2K, ces longueurs coûtent 0,80 $, 1,60 $ et 2,40 $. La génération par lots multiplie le coût basé sur la durée par le nombre de sorties.

En quoi le MiniMax H3 est-il différent du Hailuo 02 ?

MiniMax décrit Hailuo 02 comme se concentrant sur l'architecture, les données et l'échelle, tandis que MiniMax H3 se concentre sur la généralisation des tâches et des modalités. H3 ajoute également un audio stéréo natif et une sortie sélectionnable 768p ou 2K à son flux de travail texte-vidéo.

Comment passer du test du MiniMax H3 dans le navigateur à l'intégration API ?

Testez l'invite, le rapport hauteur/largeur, la résolution et la durée dans RunComfy. Appelez ensuite le même modèle Text-to-Video MiniMax H3 via l'API avec « invite » (obligatoire), « aspect_ratio », « résolution » ( « 768p » ou « 2k ») et « durée » (4-15). Ce flux de travail ne comporte aucun champ de téléchargement de média.

Suivez-nous
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • État du système
  • affilié
Modèles vidéo
  • MiniMax H3 Open
  • FLUX 3 Image to Video
  • MiniMax H3 Open Image to Video
  • Wan 2.6 Flash
  • Happy Horse 1.1 reference to video
  • Seedance 1.5 Pro Text to Video
  • Voir tous les modèles →
Modèles d'images
  • Seedream 5.0 Pro Image Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2
  • Qwen Image Edit 2511 LoRA
  • Voir tous les modèles →
Légal
  • Conditions d'utilisation
  • Politique de confidentialité
  • Politique relative aux cookies
RunComfy
Droits d'auteur 2026 RunComfy. Tous droits réservés.

RunComfy est la première ComfyUI plateforme, offrant des ComfyUI en ligne environnement et services, ainsi que des workflows ComfyUI proposant des visuels époustouflants. RunComfy propose également AI Models, permettant aux artistes d'utiliser les derniers outils d'IA pour créer des œuvres d'art incroyables.

Exemples de MiniMax H3

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...