logo
RunComfy
  • ComfyUI
  • 訓練器新
  • 模型
  • API
  • 定價
discord logo
模型
探索
所有模型
資源庫
生成記錄
模型 API
API 文檔
API 金鑰
帳戶
使用情況

MiniMax H3:帶有立體聲音訊的 768p 和 2K 文字轉視頻 |運行舒適 | Models and API | RunComfy

minimax/minimax-h3/text-to-video

MiniMax H3 文字轉影片可將書面提示轉換為具有原生立體聲的 4-15 秒 768p 或 2K 影片。使用單獨的 H3 工作流程來引用圖像、視訊或音訊。

產生影片的寬高比。
產生影片的解析度。選擇 768p 以獲得更快/更便宜的草稿,或選擇 2k 以獲得更高的細節。
產生的影片的長度(以秒為單位)(4–15)。
Idle
The rate is $0.10 per second for 768p, and $0.14 per second for 2k.

MiniMax H3 影片建立簡介

MiniMax H3 是 MiniMax 的通用多模式型號系列。此文字轉視頻頁面可將書面提示轉換為 24 FPS 的 4-15 秒視頻,並在圖片旁邊生下生動。選擇 768p 或 2K 解析度。在更廣泛的家族中,單獨的 H3 工作流程使用圖像、視訊和音訊作為身份、動作、相機、語音、聲音或編輯的參考;這些媒體輸入在此頁面上不可用。

Why Choose MiniMax H3#


MiniMax H3 is MiniMax's multimodal model family for generating and editing images, video, and audio through natural-language instructions. This page provides MiniMax H3 Text-to-Video: turn one written prompt into a 4–15-second video at 24 FPS, at either 768p or 2K, with native stereo audio generated with the picture. It accepts text only; use the linked H3 workflows when you need image, video, or audio references.


MiniMax H3 advantageWhat it means for you
Selectable 768p or 2K with native stereo audioDraft faster at 768p, or generate detailed visuals, dialogue, effects, ambience, and music in one pass at 2K—reducing separate upscaling, sound-design, and synchronization stages.
Language-directed Omni-ReferenceAcross MiniMax H3 workflows, assign images, video, and audio different roles—such as identity, product appearance, motion, camera style, voice, or music—so several sources can guide one coherent result.
V2V transfer and targeted editingMiniMax H3 can carry over motion or camera language and revise selected visual or audio elements while preserving the rest, giving production teams a path beyond regenerating a shot from scratch.
Delivery-ready timing and framingGenerate 4–15 seconds at 24 FPS in six aspect ratios, so hooks, product reveals, and multi-beat scenes fit cinematic, web, feed, portrait, or vertical placements with less recutting and recropping.

Best Use Cases#


  • Advertising and product launches: Use MiniMax H3 for hero shots, product reveals, campaign concepts, and short commercials where camera direction and sound should be designed together.
  • Social media content: Use MiniMax H3 to create 9:16 Shorts, Reels, and Stories or 1:1 feed assets with a clear opening hook and platform-ready framing.
  • Film and commercial previsualization: Use MiniMax H3 to test a camera move, lighting plan, action beat, scene transition, or temporary soundtrack before committing to production.
  • Brand and stylized storytelling: Use MiniMax H3 to explore title sequences, animated posters, character moments, and graphic looks when a written idea needs a polished audiovisual treatment.

How It Works#


  1. Write the shot: Tell MiniMax H3 what appears, what changes over time, how the camera moves, how the scene should look, and what should be heard.
  2. Choose the delivery format: Select an aspect ratio, a 4–15-second duration, and either 768p or 2k resolution.
  3. Generate and refine: The model creates video and stereo audio together. Review motion, text, faces, lip sync, and sound timing, then change one instruction at a time.

Parameters#


The first table lists the controls exposed by the MiniMax H3 Text-to-Video tool on this page.


ParameterRequiredTypeDefaultRange / OptionsHow to choose
prompt*Yes (*)StringExample prompt1–7,000 charactersGive MiniMax H3 a structured brief covering the subject, one main action, camera, setting and lighting, visual style, audio, and intended ending. Clear structure matters more than using the full limit.
aspect_ratioNoString16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Match MiniMax H3 output to the destination: 21:9 for ultra-wide cinematic shots, 16:9 for general video and ads, 4:3 for editorial or retro framing, 1:1 for feeds, 3:4 for portrait products, and 9:16 for vertical social video.
resolutionNoString768p768p, 2kUse 768p for cheaper drafts and faster iteration; use 2k when the shot needs higher detail for large placements or crops.
durationNoInteger54–15 seconds, in 1-second stepsUse 4–7-second MiniMax H3 clips for one action or fast iteration, 8–10 seconds for a simple change, and 11–15 seconds only when the prompt defines a clear beginning, development, and ending.

  • Required field.

The following table summarizes the wider MiniMax H3 model family. Inputs described for first/last-frame and Omni-Reference modes are available through separate workflows, not as controls on this Text-to-Video page.


Core dimensionMiniMax H3
ModelMiniMax-H3
Output durationMiniMax H3 outputs 4–15 seconds.
Output aspect ratioFirst/last-frame mode: follows the original aspect ratio of the input image.<br>Text-to-Video mode: follows the user-selected 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 ratio.<br>Omni-Reference mode: uses one of those six ratios or Auto, which lets MiniMax H3 choose the output ratio.
ResolutionMiniMax H3 supports two tiers. 768p: for aspect ratios from 16:9 through 9:16, the short side is 768 pixels; for wider output, total resolution is about 1 MP—for example, 21:9 is 1536×672.<br>1440p / 2K: for aspect ratios from 16:9 through 9:16, the short side is 1440 pixels; for wider output, total resolution is about 3.7 MP—for example, 21:9 is 2976×1248.
Output frame rateMiniMax H3 outputs at 24 FPS.
Output audioEvery MiniMax H3 result includes native stereo audio.
First/last-frame inputImages: 0, 1, or 2; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5). With no image input, MiniMax H3 runs in Text-to-Video mode—the tool provided on this page.
Omni-Reference inputImages: up to 9; width and height each 256–5760 pixels.<br>Videos: up to 3; each 2–15 seconds; combined video duration up to 15 seconds; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5).<br>Audio: up to 3 clips; each 2–15 seconds; combined audio duration up to 15 seconds. Audio must accompany an image or video and cannot be the only reference.<br>Mixed input: up to 12 files total. With no image, video, or audio input, the workflow becomes Text-to-Video.
Supported input formatsMiniMax H3 accepts video: H.264/AVC or H.265/HEVC; embedded audio: AAC or MP3.<br>Image: JPG, JPEG, PNG, WEBP, HEIC, or HEIF.<br>Audio: WAV or MP3.
Input size limitsEach video: 50 MB; each image: 30 MB; each audio file: 15 MB. There is no separate combined-media limit beyond the per-file limits, but the API request body is limited to 64 MB; URL-based media input is recommended.
Prompt limitThe MiniMax H3 product guide allows up to 7,000 characters. This Text-to-Video deployment currently accepts 1–7,000 characters, as shown in the page-control table above.

Pricing#


MiniMax H3 Text-to-Video pricing depends on resolution and duration:


ResolutionPrice per second5s10s15s
768p$0.10$0.50$1.00$1.50
2K$0.14$0.70$1.40$2.10

For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.


Prompting Tips & Examples#


Use this reusable MiniMax H3 prompt structure:


[subject + defining details] + [one action over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [dialogue / effects / ambience / music] + [ending frame or constraint]


  • Separate motion: Describe subject movement and camera movement independently.
  • Show sequence: Use “begins with,” “then,” and “ends on” when the clip has more than one beat.
  • Name sound sources: Tell MiniMax H3 who speaks, what makes each sound, the ambience, and whether music should lead or stay subtle.
  • Limit competing ideas: One main subject, one primary action, and one coherent style are easier for MiniMax H3 to follow.

Weak prompt


> A luxury watch ad, cinematic and dynamic, with music.


Improved MiniMax H3 prompt


> A brushed-steel automatic watch rests on black volcanic stone. A narrow studio light sweeps across the sapphire crystal as the second hand moves and condensation beads on the case. Begin with an extreme macro, then make a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Audio: soft mechanical ticking and one low cinematic pulse; no dialogue.


The improved version gives MiniMax H3 an identifiable subject, timed action, separate camera direction, lighting, finish, sound sources, and a final frame that can be reviewed.


How MiniMax H3 Compares#


Use this MiniMax H3 comparison as a model-family guide; maximum resolution and duration may not be available together, and RunComfy workflows can expose different inputs, resolution tiers, and audio controls.


ModelResolutionMax durationAudioStandout
MiniMax H3768p or 2K15sNative stereo (voice, SFX, music)Relates text, image, video, and audio references through language for V2V motion transfer and production editing; public weights are planned but have not yet been released.
Hailuo 021080p~10sNoneStrong prompt adherence and physics-focused motion suit gymnastics, dance, product movement, and other silent action shots that will receive audio later in production.
Kling 3.0Up to 4K15sNative audio with lip-syncCoordinates multi-shot camera changes with multilingual, speaker-assigned dialogue and lip-sync—useful for scripted ads, storyboards, and character-led scenes.
Seedance 2.01080p15sJoint audio and videoCombines dense image, video, and audio references with joint audiovisual generation and precise lip-sync for identity-sensitive ads, branded stories, and reference-heavy edits.
Veo 3.14K8sNative dialogue and effectsPairs prompted dialogue and effects with first/last-frame, reference-image, and scene-extension controls—suited to cinematic transitions and assembled sequences.

For Hailuo 02, the listed maxima are mode-specific: 1080p output is limited to 6 seconds, while 10-second output is available at 512p or 768p.


For MiniMax H3, Omni-Reference and V2V capabilities belong to sibling workflows; the Text-to-Video tool on this page remains prompt-only.


What sets MiniMax H3 apart is the combination: one general-purpose model family that relates text, image, video, and audio, generates native stereo sound, and offers both 768p ($0.10/s) and 2K ($0.14/s). Choose MiniMax H3 when you want to start from a text brief and keep a path to reference-guided creation or editing in sibling workflows. Based on publicly available information, run the same brief through each model before committing a pipeline.


More Models to Try#


If MiniMax H3 is not the right starting point for a project, compare these focused workflows on RunComfy:


  • MiniMax H3 Image-to-Video: Start from an image and optionally guide the final frame when composition, identity, or product appearance must stay anchored.
  • MiniMax H3 Reference-to-Video: Combine images with optional video and audio when motion, camera style, voice, music, or other source details must carry over.
  • Hailuo 02 Image-to-Video: Try an earlier MiniMax generation when you want a focused image-animation workflow.
  • Hailuo 2.3 Pro: Animate a still image in 1080p with stable, physics-aware motion, nuanced facial expression, and natural lighting—well suited to polished product, portrait, and character shots.
  • Kling 3.0: Animate a start image with an optional end frame, element references, timed shot prompts, and synchronized audio for controlled brand and character sequences.
  • Seedance 2.0 Pro: Mix up to nine images, three videos, and three audio clips to keep identity, camera language, and synchronized speech, effects, and music aligned across 4–15-second clips.

Official Resources#


  • MiniMax H3 launch post
  • MiniMax official website

相關型號

sora-2/image-to-video

將靜態圖片轉化為細膩動態影片,享受真實音畫同步的創作體驗。

seedance-2.0-mini/text-to-video

快速、低成本的多鏡頭 AI 影片模型,原生支援音訊和參考素材。

creatify/lipsync

將音訊軌道與影片同步,並設定是否循環影片。

one-to-all-animation/14b

將驅動影片中的動作遷移到參考角色圖片,並可調整提示詞、解析度、推理步數、圖片引導與姿態引導。

flux-3/image-to-video

將起始靜幀動畫化為可選原生音訊的短影片

kling-1-6/pro/image-to-video

精準提示理解、自然畫面運動與高畫質影像呈現,創造逼真AI影片。

常見問題

MiniMax H3 現在可以在 RunComfy 上使用嗎?

是的。 MiniMax H3 文字轉影片可在瀏覽器中使用,並可使用您的帳戶積分透過 RunComfy API 使用。

什麼是 MiniMax H3?

MiniMax H3 是 MiniMax 的通用多模式產生和編輯模型系列。其更廣泛的任務設計涵蓋文字、圖像、視訊和音頻,而各個工作流程則公開不同的輸入。此頁面提供僅提示的文字轉影片工作流程。

MiniMax H3 可以用來做什麼?

MiniMax 將 H3 定位於廣告、品牌、電子商務、電影、標題設計、動畫海報、短片故事、產品和 UI 概念、遊戲、虛擬角色和風格化動畫。目前的文字到視訊工作流程最適合可以透過書面提示指導的簡短概念。

MiniMax H3 文字轉影片支援什麼解析度和剪輯長度?

在此頁面上,MiniMax H3 支援「768p」和「2k」解析度。選擇 4 到 15 秒之間的任何整秒持續時間;預設值為 5 秒。對於更便宜的草稿,請使用 768p;當您需要更高的細節時,請使用 2K。

MiniMax H3 可以產生音訊嗎?

是的。此文字到視訊工作流程會隨視訊生成本機立體聲音訊。描述提示中的對話、氛圍、音效或音樂;沒有單獨的音訊切換。結果可能會有所不同,因此請檢查口型同步和音訊時序。

我可以在此 MiniMax H3 頁面上使用圖像、視訊或音訊參考嗎?

不可以。此頁面是僅提示的 MiniMax H3 文字轉視訊工作流程,其 API 僅接受「提示」、「寬高比」、「解析度」和「持續時間」。對於來源媒體,請使用單獨的 H3 影像到影片或參考到視訊工作流程。

多模式環境對於 MiniMax H3 車型系列意味著什麼?

在支援參考的 H3 工作流程中,自然語言可以為文字、圖像、視訊和音訊分配不同的角色。例如,一個來源可能定義攝影機運動,另一個來源定義角色,另一個來源定義聲音。這是模型系列功能,而不是目前頁面上的媒體上傳功能。

我需要自行託管 MiniMax H3 嗎?

不需要。 RunComfy 透過瀏覽器和 HTTP API 提供 MiniMax H3,因此您無需自行託管或擴充模型。 MiniMax 在 7 月 31 日發布的帖子中描述了一項有條件的發布模型權重的計劃;在規劃自託管之前驗證當前的可用性和許可條款。

MiniMax H3 文字轉影片的費用是多少?

MiniMax H3 文字轉影片在 768p 下每產生一秒鐘的成本為 0.10 美元,在 2K 下每產生一秒的成本為 0.14 美元。在 768p 下,5 秒影片的成本為 0.50 美元,10 秒影片的成本為 1.00 美元,15 秒影片的成本為 1.50 美元。在 2K 時,這些長度的價格為 0.70 美元、1.40 美元和 2.10 美元。批量產生將基於持續時間的成本乘以輸出數量。

MiniMax H3 與海螺 02 有何不同?

MiniMax 將 Hailuo 02 描述為專注於架構、資料和規模,而 MiniMax H3 則專注於泛化任務和模式。 H3 還在其文字轉視訊工作流程中加入了原生立體聲音訊和可選擇的 768p 或 2K 輸出。

如何從在瀏覽器中測試 MiniMax H3 到 API 整合?

在 RunComfy 中測試提示、寬高比、解析度和持續時間。然後透過 API 使用「提示」(必需)、「寬高比」、「解析度」(「768p」或「2k」)和「持續時間」(4-15)呼叫相同的 MiniMax H3 文字到視訊範本。此工作流程沒有媒體上傳欄位。

關注我們
  • 領英
  • Facebook
  • Instagram
  • Twitter
支持
  • Discord
  • 電子郵件
  • 系統狀態
  • 附屬
視頻模型
  • MiniMax H3 Open
  • FLUX 3 Image to Video
  • MiniMax H3 Open Image to Video
  • Wan 2.6 Flash
  • Happy Horse 1.1 reference to video
  • Seedance 1.5 Pro Text to Video
  • 查看所有模型 →
影像模型
  • Seedream 5.0 Pro Image Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2
  • Qwen Image Edit 2511 LoRA
  • 查看所有模型 →
法律
  • 服務條款
  • 隱私政策
  • Cookie 政策
RunComfy
版權 2026 RunComfy. 保留所有權利。

RunComfy 是首選的 ComfyUI 平台,提供 ComfyUI 在線 環境和服務,以及 ComfyUI 工作流程 具有驚豔的視覺效果。 RunComfy還提供 AI Models, 幫助藝術家利用最新的AI工具創作出令人驚艷的藝術作品。

MiniMax H3 範例

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...