根据文本标签生成最长 4 分钟、带人声与歌词的歌曲。
MiniMax H3 is MiniMax's multimodal model family for generating and editing images, video, and audio through natural-language instructions. This page provides MiniMax H3 Text-to-Video: turn one written prompt into a 4–15-second video at 24 FPS, at either 768p or 2K, with native stereo audio generated with the picture. It accepts text only; use the linked H3 workflows when you need image, video, or audio references.
| MiniMax H3 advantage | What it means for you |
|---|---|
| Selectable 768p or 2K with native stereo audio | Draft faster at 768p, or generate detailed visuals, dialogue, effects, ambience, and music in one pass at 2K—reducing separate upscaling, sound-design, and synchronization stages. |
| Language-directed Omni-Reference | Across MiniMax H3 workflows, assign images, video, and audio different roles—such as identity, product appearance, motion, camera style, voice, or music—so several sources can guide one coherent result. |
| V2V transfer and targeted editing | MiniMax H3 can carry over motion or camera language and revise selected visual or audio elements while preserving the rest, giving production teams a path beyond regenerating a shot from scratch. |
| Delivery-ready timing and framing | Generate 4–15 seconds at 24 FPS in six aspect ratios, so hooks, product reveals, and multi-beat scenes fit cinematic, web, feed, portrait, or vertical placements with less recutting and recropping. |
9:16 Shorts, Reels, and Stories or 1:1 feed assets with a clear opening hook and platform-ready framing.768p or 2k resolution.The first table lists the controls exposed by the MiniMax H3 Text-to-Video tool on this page.
| Parameter | Required | Type | Default | Range / Options | How to choose |
|---|---|---|---|---|---|
prompt* | Yes (*) | String | Example prompt | 1–4,000 characters | Give MiniMax H3 a structured brief covering the subject, one main action, camera, setting and lighting, visual style, audio, and intended ending. Clear structure matters more than using the full limit. |
aspect_ratio | No | String | 16:9 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Match MiniMax H3 output to the destination: 21:9 for ultra-wide cinematic shots, 16:9 for general video and ads, 4:3 for editorial or retro framing, 1:1 for feeds, 3:4 for portrait products, and 9:16 for vertical social video. |
resolution | No | String | 768p | 768p, 2k | Use 768p for cheaper drafts and faster iteration; use 2k when the shot needs higher detail for large placements or crops. |
duration | No | Integer | 5 | 4–15 seconds, in 1-second steps | Use 4–7-second MiniMax H3 clips for one action or fast iteration, 8–10 seconds for a simple change, and 11–15 seconds only when the prompt defines a clear beginning, development, and ending. |
The following table summarizes the wider MiniMax H3 model family. Inputs described for first/last-frame and Omni-Reference modes are available through separate workflows, not as controls on this Text-to-Video page.
| Core dimension | MiniMax H3 |
|---|---|
| Model | MiniMax-H3 |
| Output duration | MiniMax H3 outputs 4–15 seconds. |
| Output aspect ratio | First/last-frame mode: follows the original aspect ratio of the input image.<br>Text-to-Video mode: follows the user-selected 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 ratio.<br>Omni-Reference mode: uses one of those six ratios or Auto, which lets MiniMax H3 choose the output ratio. |
| Resolution | MiniMax H3 supports two tiers. 768p: for aspect ratios from 16:9 through 9:16, the short side is 768 pixels; for wider output, total resolution is about 1 MP—for example, 21:9 is 1536×672.<br>1440p / 2K: for aspect ratios from 16:9 through 9:16, the short side is 1440 pixels; for wider output, total resolution is about 3.7 MP—for example, 21:9 is 2976×1248. |
| Output frame rate | MiniMax H3 outputs at 24 FPS. |
| Output audio | Every MiniMax H3 result includes native stereo audio. |
| First/last-frame input | Images: 0, 1, or 2; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5). With no image input, MiniMax H3 runs in Text-to-Video mode—the tool provided on this page. |
| Omni-Reference input | Images: up to 9; width and height each 256–5760 pixels.<br>Videos: up to 3; each 2–15 seconds; combined video duration up to 15 seconds; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5).<br>Audio: up to 3 clips; each 2–15 seconds; combined audio duration up to 15 seconds. Audio must accompany an image or video and cannot be the only reference.<br>Mixed input: up to 12 files total. With no image, video, or audio input, the workflow becomes Text-to-Video. |
| Supported input formats | MiniMax H3 accepts video: H.264/AVC or H.265/HEVC; embedded audio: AAC or MP3.<br>Image: JPG, JPEG, PNG, WEBP, HEIC, or HEIF.<br>Audio: WAV or MP3. |
| Input size limits | Each video: 50 MB; each image: 30 MB; each audio file: 15 MB. There is no separate combined-media limit beyond the per-file limits, but the API request body is limited to 64 MB; URL-based media input is recommended. |
| Prompt limit | The MiniMax H3 product guide allows up to 7,000 characters. This Text-to-Video deployment currently accepts 1–4,000 characters, as shown in the page-control table above. |
MiniMax H3 Text-to-Video pricing depends on resolution and duration:
| Resolution | Price per second | 5s | 10s | 15s |
|---|---|---|---|---|
| 768p | $0.11 | $0.55 | $1.10 | $1.65 |
| 2K | $0.16 | $0.80 | $1.60 | $2.40 |
For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.
Use this reusable MiniMax H3 prompt structure:
[subject + defining details] + [one action over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [dialogue / effects / ambience / music] + [ending frame or constraint]
Weak prompt
> A luxury watch ad, cinematic and dynamic, with music.
Improved MiniMax H3 prompt
> A brushed-steel automatic watch rests on black volcanic stone. A narrow studio light sweeps across the sapphire crystal as the second hand moves and condensation beads on the case. Begin with an extreme macro, then make a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Audio: soft mechanical ticking and one low cinematic pulse; no dialogue.
The improved version gives MiniMax H3 an identifiable subject, timed action, separate camera direction, lighting, finish, sound sources, and a final frame that can be reviewed.
Use this MiniMax H3 comparison as a model-family guide; maximum resolution and duration may not be available together, and RunComfy workflows can expose different inputs, resolution tiers, and audio controls.
| Model | Resolution | Max duration | Audio | Standout |
|---|---|---|---|---|
| MiniMax H3 | 768p or 2K | 15s | Native stereo (voice, SFX, music) | Relates text, image, video, and audio references through language for V2V motion transfer and production editing; public weights are planned but have not yet been released. |
| Hailuo 02 | 1080p | ~10s | None | Strong prompt adherence and physics-focused motion suit gymnastics, dance, product movement, and other silent action shots that will receive audio later in production. |
| Kling 3.0 | Up to 4K | 15s | Native audio with lip-sync | Coordinates multi-shot camera changes with multilingual, speaker-assigned dialogue and lip-sync—useful for scripted ads, storyboards, and character-led scenes. |
| Seedance 2.0 | 1080p | 15s | Joint audio and video | Combines dense image, video, and audio references with joint audiovisual generation and precise lip-sync for identity-sensitive ads, branded stories, and reference-heavy edits. |
| Veo 3.1 | 4K | 8s | Native dialogue and effects | Pairs prompted dialogue and effects with first/last-frame, reference-image, and scene-extension controls—suited to cinematic transitions and assembled sequences. |
For Hailuo 02, the listed maxima are mode-specific: 1080p output is limited to 6 seconds, while 10-second output is available at 512p or 768p.
For MiniMax H3, Omni-Reference and V2V capabilities belong to sibling workflows; the Text-to-Video tool on this page remains prompt-only.
What sets MiniMax H3 apart is the combination: one general-purpose model family that relates text, image, video, and audio, generates native stereo sound, and offers both 768p ($0.11/s) and 2K ($0.16/s). Choose MiniMax H3 when you want to start from a text brief and keep a path to reference-guided creation or editing in sibling workflows. Based on publicly available information, run the same brief through each model before committing a pipeline.
If MiniMax H3 is not the right starting point for a project, compare these focused workflows on RunComfy:
根据文本标签生成最长 4 分钟、带人声与歌词的歌曲。
根据驱动视频中的动作,让参考图片动起来。
从图像、视频和音频参考生成 768p 或 2K 视频
Kling V3.0 系列中具有最高视觉保真度的优质电影文本转视频。
使用一张人像和一段音频生成口播数字人视频,并可通过提示词引导动作与表达方式。
使用 Sora 2 Pro 根据文本提示词生成视频,并选择输出尺寸和时长。
是的。 MiniMax H3 文本到视频可在浏览器中使用,并可使用您的帐户积分通过 RunComfy API 使用。
MiniMax H3 是 MiniMax 的通用多模式生成和编辑模型系列。其更广泛的任务设计涵盖文本、图像、视频和音频,而各个工作流程则公开不同的输入。此页面提供仅提示的文本转视频工作流程。
MiniMax 将 H3 定位于广告、品牌、电子商务、电影、标题设计、动画海报、短片故事、产品和 UI 概念、游戏、虚拟角色和风格化动画。当前的文本到视频工作流程最适合可以通过书面提示指导的简短概念。
在此页面上,MiniMax H3 支持“768p”和“2k”分辨率。选择 4 到 15 秒之间的任何整秒持续时间;默认值为 5 秒。对于更便宜的草稿,请使用 768p;当您需要更高的细节时,请使用 2K。
是的。此文本到视频工作流程会随视频生成本机立体声音频。描述提示中的对话、氛围、音效或音乐;没有单独的音频切换。结果可能会有所不同,因此请检查口型同步和音频时序。
不可以。此页面是仅提示的 MiniMax H3 文本转视频工作流程,其 API 仅接受“提示”、“宽高比”、“分辨率”和“持续时间”。对于源媒体,请使用单独的 H3 图像到视频或参考到视频工作流程。
在支持参考的 H3 工作流程中,自然语言可以为文本、图像、视频和音频分配不同的角色。例如,一个源可能定义摄像机运动,另一个源定义角色,另一个源定义声音。这是模型系列功能,而不是当前页面上的媒体上传功能。
不需要。RunComfy 通过浏览器和 HTTP API 提供 MiniMax H3,因此您无需自行托管或扩展模型。 MiniMax 在 7 月 31 日发布的帖子中描述了一项有条件的发布模型权重的计划;在规划自托管之前验证当前的可用性和许可条款。
MiniMax H3 文本转视频在 768p 下每生成一秒的成本为 0.11 美元,在 2K 下每生成一秒的成本为 0.16 美元。在 768p 下,5 秒视频的成本为 0.55 美元,10 秒视频的成本为 1.10 美元,15 秒视频的成本为 1.65 美元。在 2K 时,这些长度的价格为 0.80 美元、1.60 美元和 2.40 美元。批量生成将基于持续时间的成本乘以输出数量。
MiniMax 将 Hailuo 02 描述为专注于架构、数据和规模,而 MiniMax H3 则专注于泛化任务和模式。 H3 还在其文本转视频工作流程中添加了原生立体声音频和可选择的 768p 或 2K 输出。
在 RunComfy 中测试提示、宽高比、分辨率和持续时间。然后通过 API 使用“提示”(必需)、“宽高比”、“分辨率”(“768p”或“2k”)和“持续时间”(4-15)调用相同的 MiniMax H3 文本到视频模板。此工作流程没有媒体上传字段。
RunComfy 是首选的 ComfyUI 平台,提供 ComfyUI 在线 环境和服务,以及 ComfyUI 工作流 具有惊艳的视觉效果。 RunComfy还提供 AI Models, 帮助艺术家利用最新的AI工具创作出令人惊叹的艺术作品。














