logo
RunComfy
  • ComfyUI
  • 트레이너신규
  • 모델
  • API
  • 가격
discord logo
모델
탐색
모든 모델
라이브러리
생성 기록
모델 API
API 문서
API 키
계정
사용량

MiniMax H3: 스테레오 오디오를 갖춘 768p 및 2K 텍스트-비디오 | 달리다편안하다 | Models and API | RunComfy

minimax/minimax-h3/text-to-video

MiniMax H3 텍스트-비디오는 서면 메시지를 기본 스테레오 사운드가 포함된 4~15초 길이의 768p 또는 2K 비디오로 변환합니다. 이미지, 비디오, 오디오 참조에는 별도의 H3 워크플로를 사용하세요.

생성된 비디오의 화면 비율입니다.
생성된 비디오의 해상도입니다. 더 빠르고 저렴한 초안을 원하면 768p를 선택하고, 더 자세한 내용을 원하면 2k를 선택하세요.
생성된 비디오의 길이(초)입니다(4~15).
Idle
The rate is $0.11 per second for 768p, and $0.16 per second for 2k.

MiniMax H3 비디오 제작 소개

MiniMax H3는 MiniMax의 범용 다중 모드 모델 제품군입니다. 이 텍스트-비디오 페이지는 서면 메시지를 그림과 함께 생성된 스테레오 사운드와 함께 24FPS의 4~15초 비디오로 변환합니다. 768p 또는 2K 해상도를 선택하세요. 광범위한 제품군에 걸쳐 별도의 H3 워크플로우는 이미지, 비디오 및 오디오를 ID, 동작, 카메라, 음성, 사운드 또는 편집에 대한 참조로 사용합니다. 해당 미디어 입력은 이 페이지에서 사용할 수 없습니다.

Why Choose MiniMax H3#


MiniMax H3 is MiniMax's multimodal model family for generating and editing images, video, and audio through natural-language instructions. This page provides MiniMax H3 Text-to-Video: turn one written prompt into a 4–15-second video at 24 FPS, at either 768p or 2K, with native stereo audio generated with the picture. It accepts text only; use the linked H3 workflows when you need image, video, or audio references.


MiniMax H3 advantageWhat it means for you
Selectable 768p or 2K with native stereo audioDraft faster at 768p, or generate detailed visuals, dialogue, effects, ambience, and music in one pass at 2K—reducing separate upscaling, sound-design, and synchronization stages.
Language-directed Omni-ReferenceAcross MiniMax H3 workflows, assign images, video, and audio different roles—such as identity, product appearance, motion, camera style, voice, or music—so several sources can guide one coherent result.
V2V transfer and targeted editingMiniMax H3 can carry over motion or camera language and revise selected visual or audio elements while preserving the rest, giving production teams a path beyond regenerating a shot from scratch.
Delivery-ready timing and framingGenerate 4–15 seconds at 24 FPS in six aspect ratios, so hooks, product reveals, and multi-beat scenes fit cinematic, web, feed, portrait, or vertical placements with less recutting and recropping.

Best Use Cases#


  • Advertising and product launches: Use MiniMax H3 for hero shots, product reveals, campaign concepts, and short commercials where camera direction and sound should be designed together.
  • Social media content: Use MiniMax H3 to create 9:16 Shorts, Reels, and Stories or 1:1 feed assets with a clear opening hook and platform-ready framing.
  • Film and commercial previsualization: Use MiniMax H3 to test a camera move, lighting plan, action beat, scene transition, or temporary soundtrack before committing to production.
  • Brand and stylized storytelling: Use MiniMax H3 to explore title sequences, animated posters, character moments, and graphic looks when a written idea needs a polished audiovisual treatment.

How It Works#


  1. Write the shot: Tell MiniMax H3 what appears, what changes over time, how the camera moves, how the scene should look, and what should be heard.
  2. Choose the delivery format: Select an aspect ratio, a 4–15-second duration, and either 768p or 2k resolution.
  3. Generate and refine: The model creates video and stereo audio together. Review motion, text, faces, lip sync, and sound timing, then change one instruction at a time.

Parameters#


The first table lists the controls exposed by the MiniMax H3 Text-to-Video tool on this page.


ParameterRequiredTypeDefaultRange / OptionsHow to choose
prompt*Yes (*)StringExample prompt1–4,000 charactersGive MiniMax H3 a structured brief covering the subject, one main action, camera, setting and lighting, visual style, audio, and intended ending. Clear structure matters more than using the full limit.
aspect_ratioNoString16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Match MiniMax H3 output to the destination: 21:9 for ultra-wide cinematic shots, 16:9 for general video and ads, 4:3 for editorial or retro framing, 1:1 for feeds, 3:4 for portrait products, and 9:16 for vertical social video.
resolutionNoString768p768p, 2kUse 768p for cheaper drafts and faster iteration; use 2k when the shot needs higher detail for large placements or crops.
durationNoInteger54–15 seconds, in 1-second stepsUse 4–7-second MiniMax H3 clips for one action or fast iteration, 8–10 seconds for a simple change, and 11–15 seconds only when the prompt defines a clear beginning, development, and ending.

  • Required field.

The following table summarizes the wider MiniMax H3 model family. Inputs described for first/last-frame and Omni-Reference modes are available through separate workflows, not as controls on this Text-to-Video page.


Core dimensionMiniMax H3
ModelMiniMax-H3
Output durationMiniMax H3 outputs 4–15 seconds.
Output aspect ratioFirst/last-frame mode: follows the original aspect ratio of the input image.<br>Text-to-Video mode: follows the user-selected 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 ratio.<br>Omni-Reference mode: uses one of those six ratios or Auto, which lets MiniMax H3 choose the output ratio.
ResolutionMiniMax H3 supports two tiers. 768p: for aspect ratios from 16:9 through 9:16, the short side is 768 pixels; for wider output, total resolution is about 1 MP—for example, 21:9 is 1536×672.<br>1440p / 2K: for aspect ratios from 16:9 through 9:16, the short side is 1440 pixels; for wider output, total resolution is about 3.7 MP—for example, 21:9 is 2976×1248.
Output frame rateMiniMax H3 outputs at 24 FPS.
Output audioEvery MiniMax H3 result includes native stereo audio.
First/last-frame inputImages: 0, 1, or 2; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5). With no image input, MiniMax H3 runs in Text-to-Video mode—the tool provided on this page.
Omni-Reference inputImages: up to 9; width and height each 256–5760 pixels.<br>Videos: up to 3; each 2–15 seconds; combined video duration up to 15 seconds; width and height each 256–5760 pixels; aspect ratio from 5:2 to 2:5 (0.4–2.5).<br>Audio: up to 3 clips; each 2–15 seconds; combined audio duration up to 15 seconds. Audio must accompany an image or video and cannot be the only reference.<br>Mixed input: up to 12 files total. With no image, video, or audio input, the workflow becomes Text-to-Video.
Supported input formatsMiniMax H3 accepts video: H.264/AVC or H.265/HEVC; embedded audio: AAC or MP3.<br>Image: JPG, JPEG, PNG, WEBP, HEIC, or HEIF.<br>Audio: WAV or MP3.
Input size limitsEach video: 50 MB; each image: 30 MB; each audio file: 15 MB. There is no separate combined-media limit beyond the per-file limits, but the API request body is limited to 64 MB; URL-based media input is recommended.
Prompt limitThe MiniMax H3 product guide allows up to 7,000 characters. This Text-to-Video deployment currently accepts 1–4,000 characters, as shown in the page-control table above.

Pricing#


MiniMax H3 Text-to-Video pricing depends on resolution and duration:


ResolutionPrice per second5s10s15s
768p$0.11$0.55$1.10$1.65
2K$0.16$0.80$1.60$2.40

For batches of 1–4 outputs, calculate the total as duration × per-second rate × output count.


Prompting Tips & Examples#


Use this reusable MiniMax H3 prompt structure:


[subject + defining details] + [one action over time] + [camera framing and movement] + [setting + lighting] + [visual treatment] + [dialogue / effects / ambience / music] + [ending frame or constraint]


  • Separate motion: Describe subject movement and camera movement independently.
  • Show sequence: Use “begins with,” “then,” and “ends on” when the clip has more than one beat.
  • Name sound sources: Tell MiniMax H3 who speaks, what makes each sound, the ambience, and whether music should lead or stay subtle.
  • Limit competing ideas: One main subject, one primary action, and one coherent style are easier for MiniMax H3 to follow.

Weak prompt


> A luxury watch ad, cinematic and dynamic, with music.


Improved MiniMax H3 prompt


> A brushed-steel automatic watch rests on black volcanic stone. A narrow studio light sweeps across the sapphire crystal as the second hand moves and condensation beads on the case. Begin with an extreme macro, then make a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Audio: soft mechanical ticking and one low cinematic pulse; no dialogue.


The improved version gives MiniMax H3 an identifiable subject, timed action, separate camera direction, lighting, finish, sound sources, and a final frame that can be reviewed.


How MiniMax H3 Compares#


Use this MiniMax H3 comparison as a model-family guide; maximum resolution and duration may not be available together, and RunComfy workflows can expose different inputs, resolution tiers, and audio controls.


ModelResolutionMax durationAudioStandout
MiniMax H3768p or 2K15sNative stereo (voice, SFX, music)Relates text, image, video, and audio references through language for V2V motion transfer and production editing; public weights are planned but have not yet been released.
Hailuo 021080p~10sNoneStrong prompt adherence and physics-focused motion suit gymnastics, dance, product movement, and other silent action shots that will receive audio later in production.
Kling 3.0Up to 4K15sNative audio with lip-syncCoordinates multi-shot camera changes with multilingual, speaker-assigned dialogue and lip-sync—useful for scripted ads, storyboards, and character-led scenes.
Seedance 2.01080p15sJoint audio and videoCombines dense image, video, and audio references with joint audiovisual generation and precise lip-sync for identity-sensitive ads, branded stories, and reference-heavy edits.
Veo 3.14K8sNative dialogue and effectsPairs prompted dialogue and effects with first/last-frame, reference-image, and scene-extension controls—suited to cinematic transitions and assembled sequences.

For Hailuo 02, the listed maxima are mode-specific: 1080p output is limited to 6 seconds, while 10-second output is available at 512p or 768p.


For MiniMax H3, Omni-Reference and V2V capabilities belong to sibling workflows; the Text-to-Video tool on this page remains prompt-only.


What sets MiniMax H3 apart is the combination: one general-purpose model family that relates text, image, video, and audio, generates native stereo sound, and offers both 768p ($0.11/s) and 2K ($0.16/s). Choose MiniMax H3 when you want to start from a text brief and keep a path to reference-guided creation or editing in sibling workflows. Based on publicly available information, run the same brief through each model before committing a pipeline.


More Models to Try#


If MiniMax H3 is not the right starting point for a project, compare these focused workflows on RunComfy:


  • MiniMax H3 Image-to-Video: Start from an image and optionally guide the final frame when composition, identity, or product appearance must stay anchored.
  • MiniMax H3 Reference-to-Video: Combine images with optional video and audio when motion, camera style, voice, music, or other source details must carry over.
  • Hailuo 02 Image-to-Video: Try an earlier MiniMax generation when you want a focused image-animation workflow.
  • Hailuo 2.3 Pro: Animate a still image in 1080p with stable, physics-aware motion, nuanced facial expression, and natural lighting—well suited to polished product, portrait, and character shots.
  • Kling 3.0: Animate a start image with an optional end frame, element references, timed shot prompts, and synchronized audio for controlled brand and character sequences.
  • Seedance 2.0 Pro: Mix up to nine images, three videos, and three audio clips to keep identity, camera language, and synchronized speech, effects, and music aligned across 4–15-second clips.

Official Resources#


  • MiniMax H3 launch post
  • MiniMax official website

관련 모델

dreamina-3-0/image-to-video

Dreamina 3.0으로 정적인 이미지를 2K 수준의 생동감 있는 영상으로 손쉽게 변환해보세요.

kling-video-o3/4K/text-to-video

출력 1초당 $0.42로 만드는 영화 같은 4K 텍스트-투-비디오.

veo-3-1/fast/image-to-video

이미지로부터 영화 같은 영상을 빠르고 자연스럽게 제작하는 Veo 3.1 Fast.

hunyuan-video-v1.5/text-to-video

필수 프롬프트를 바탕으로 5초 또는 8초 동영상을 생성합니다. 선택형 네거티브 프롬프트, 네 가지 출력 크기와 시드를 설정할 수 있습니다.

flux-3/first-last-frame-to-video

시작·끝 프레임 사이로 선택적 네이티브 오디오 영상 생성

dreamina-3-0/pro/image-to-video

정적인 이미지를 사실적인 움직임으로 바꾸는 고화질 AI 영상 생성 도구.

자주 묻는 질문

이제 RunComfy에서 MiniMax H3를 사용할 수 있나요?

예. MiniMax H3 텍스트-비디오는 계정 크레딧을 사용하여 브라우저와 RunComfy API를 통해 사용할 수 있습니다.

미니맥스 H3란 무엇인가요?

MiniMax H3는 MiniMax의 범용 다중 모드 생성 및 편집 모델 제품군입니다. 더 넓은 작업 디자인에는 텍스트, 이미지, 비디오 및 오디오가 포함되며 개별 워크플로는 다양한 입력을 제공합니다. 이 페이지는 프롬프트 전용 텍스트-비디오 워크플로우를 제공합니다.

MiniMax H3는 어떤 용도로 사용할 수 있나요?

MiniMax는 광고, 브랜딩, 전자상거래, 영화, 타이틀 디자인, 애니메이션 포스터, 단편 스토리텔링, 제품 및 UI 컨셉, 게임, 가상 캐릭터, 스타일화된 애니메이션 분야에서 H3를 포지셔닝합니다. 현재의 텍스트-비디오 워크플로우는 서면 프롬프트에서 지시할 수 있는 짧은 개념에 가장 적합합니다.

MiniMax H3 Text-to-Video는 어떤 해상도와 클립 길이를 지원합니까?

이 페이지에서 MiniMax H3는 '768p' 및 '2k' 해상도를 지원합니다. 4초에서 15초 사이에서 전체 초 기간을 선택합니다. 기본값은 5초입니다. 저렴한 초안에는 768p를 사용하고, 더 높은 세부 묘사가 필요할 때는 2K를 사용하세요.

MiniMax H3는 오디오를 생성합니까?

예. 이 텍스트-비디오 워크플로우는 비디오와 함께 기본 스테레오 오디오를 생성합니다. 프롬프트에서 대화, 분위기, 음향 효과 또는 음악을 설명합니다. 별도의 오디오 토글이 없습니다. 결과는 다양할 수 있으므로 립싱크와 오디오 타이밍을 검토하세요.

이 MiniMax H3 페이지에서 이미지, 비디오 또는 오디오 참조를 사용할 수 있습니까?

아니요. 이 페이지는 프롬프트 전용 MiniMax H3 텍스트-비디오 워크플로이며 해당 API는 '프롬프트', '종횡비', '해상도' 및 '기간'만 허용합니다. 소스 미디어의 경우 별도의 H3 이미지-비디오 또는 참조-비디오 워크플로를 사용하세요.

MiniMax H3 모델 제품군에 있어 다중 모드 컨텍스트는 무엇을 의미합니까?

참조 지원 H3 워크플로에서 자연어는 텍스트, 이미지, 비디오 및 오디오에 다양한 역할을 할당할 수 있습니다. 예를 들어 한 소스는 카메라 움직임을 정의하고 다른 소스는 캐릭터, 다른 소스는 음성을 정의할 수 있습니다. 이는 현재 페이지의 미디어 업로드 기능이 아닌 모델 계열 기능입니다.

MiniMax H3를 자체 호스팅해야 합니까?

아니요. RunComfy는 브라우저와 HTTP API를 통해 MiniMax H3를 제공하므로 모델을 직접 호스팅하거나 확장할 필요가 없습니다. MiniMax의 7월 31일 출시 게시물에서는 모델 가중치를 출시하기 위한 조건부 계획을 설명했습니다. 자체 호스팅을 계획하기 전에 현재 가용성과 라이선스 조건을 확인하세요.

MiniMax H3 텍스트-비디오 비용은 얼마입니까?

MiniMax H3 텍스트-비디오 비용은 768p에서 생성된 초당 $0.11, 2K에서 생성된 초당 $0.16입니다. 768p에서 5초짜리 비디오 비용은 $0.55, 10초 비디오 비용은 $1.10, 15초 비디오 비용은 $1.65입니다. 2K에서 해당 길이의 비용은 $0.80, $1.60, $2.40입니다. 일괄 생성에서는 기간 기반 비용에 출력 수를 곱합니다.

MiniMax H3는 Hailuo 02와 어떻게 다릅니까?

MiniMax는 Hailuo 02가 아키텍처, 데이터 및 규모에 중점을 두고 있는 반면 MiniMax H3는 일반화 작업 및 양식에 중점을 두고 있다고 설명합니다. H3는 또한 텍스트-비디오 워크플로우에 기본 스테레오 오디오와 선택 가능한 768p 또는 2K 출력을 추가합니다.

브라우저에서 MiniMax H3 테스트를 API 통합으로 어떻게 진행하나요?

RunComfy에서 프롬프트, 종횡비, 해상도 및 지속 시간을 테스트하세요. 그런 다음 prompt(필수), aspect_ratio, solution(768p 또는 2k) 및 duration(4–15)을 사용하여 API를 통해 동일한 MiniMax H3 텍스트-비디오 템플릿을 호출합니다. 이 워크플로에는 미디어 업로드 필드가 없습니다.

팔로우하기
  • 링크드인
  • 페이스북
  • Instagram
  • 트위터
지원
  • 디스코드
  • 이메일
  • 시스템 상태
  • 제휴사
비디오 모델
  • MiniMax H3 Open
  • FLUX 3 Image to Video
  • MiniMax H3 Open Image to Video
  • Wan 2.6 Flash
  • Happy Horse 1.1 reference to video
  • Seedance 1.5 Pro Text to Video
  • 모든 모델 보기 →
이미지 모델
  • Seedream 5.0 Pro Image Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2
  • Qwen Image Edit 2511 LoRA
  • 모든 모델 보기 →
법적 고지
  • 서비스 약관
  • 개인정보 보호정책
  • 쿠키 정책
RunComfy
저작권 2026 RunComfy. All Rights Reserved.

RunComfy는 최고의 ComfyUI 플랫폼으로서 ComfyUI 온라인 환경과 서비스를 제공하며 ComfyUI 워크플로우 멋진 비주얼을 제공합니다. RunComfy는 또한 제공합니다 AI Models, 예술가들이 최신 AI 도구를 활용하여 놀라운 예술을 창조할 수 있도록 지원합니다.

MiniMax H3의 예

Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...
Video thumbnail
Loading...