必須のプロンプトから5秒または8秒の動画を生成します。任意のネガティブプロンプト、4種類の出力サイズ、シードを設定できます。
Minimax H3 Reference to Video generates a 768p, 2K clip from a text prompt and the reference media you attach. Supply at least one reference image or video, describe the scene you want, and the model works out how your references should behave on screen.
The split is simple. Images pin what things look like; videos pin how things move. Minimax H3 Reference to Video reads both against your prompt rather than treating either as a fixed template.
768p for cheaper drafts or 2k when detail needs to hold up in large placements.The live Minimax H3 Reference to Video fields on this page.
| Parameter | Required | Type | Default | Range / Options | Description |
|---|---|---|---|---|---|
| prompt * | Yes (*) | string | Sample brief | 1-4000 characters | Scene, motion, camera work, visual style. |
| reference_images * | Yes (*) | array (URLs) | Sample reference | Up to 9 | Images guiding subject, style, composition. |
| reference_videos | No | array (URLs) | None | Up to 3 | Clips guiding movement, pacing, or interaction. |
| reference_audios | No | array (URLs) | None | Up to 3 | Audio guidance; needs an image or video too. |
| aspect_ratio * | Yes (*) | string | 16:9 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Output framing. |
| duration * | Yes (*) | integer | 5 | 4-15 | Clip length in whole seconds. |
| resolution * | Yes (*) | string | 768p | 768p, 2k | Output resolution tier. |
At least one reference image or video is required.
Minimax H3 Reference to Video bills by resolution on counted seconds: the finished clip plus the combined length of any reference videos attached. Reference images and audio are not billed by duration.
| Resolution | Price per counted second |
|---|---|
| 768p | $0.10 |
| 2K | $0.16 |
| Output | Reference video | Counted seconds | Cost at 768p | Cost at 2K | Cost at |
|---|---|---|---|---|---|
| 5s | none | 5 | $0.50 | $0.80 | $0.88 |
| 10s | none | 10 | $1.00 | $1.60 | $1.76 |
| 15s | none | 15 | $1.50 | $2.40 | $2.64 |
| 10s | 5s | 15 | $1.50 | $2.40 | $2.64 |
Reference clips are measured after the job runs, so the figure shown before you submit is an estimate and the final charge reflects the real counted seconds.
1) Attach your references — Add images that define the subject and, if movement matters, clips that define it. Minimax H3 Reference to Video needs at least one of the two.
2) Write the brief — Describe the action, camera work, lighting, and mood.
3) Say how the references relate — Name which one supplies the subject and which supplies the motion. Minimax H3 Reference to Video follows an explicit relationship far better than an implied one.
4) Add reference audio (optional) — Only alongside an image or video reference in Minimax H3 Reference to Video.
5) Pick the aspect ratio — 16:9 widescreen, 9:16 vertical, 1:1 square, 21:9 wide cinematic.
6) Choose resolution and length — Start at 768p and a short duration while testing, then switch to 2k and extend toward 15 seconds once Minimax H3 Reference to Video is behaving.
7) Generate and refine — Change one reference or clause at a time so you can attribute the difference in the Minimax H3 Reference to Video output.
必須のプロンプトから5秒または8秒の動画を生成します。任意のネガティブプロンプト、4種類の出力サイズ、シードを設定できます。
音声トラックを動画に同期し、動画を繰り返すかどうかを設定します。
画像を高品質な動画へ。Wan 2.5で手軽に創造的な映像制作を実現。
画像をなめらかな映像へ。自然な動きと一貫した表現で創造力を広げるAI動画生成ツール。
テキストプロンプトから16:9動画を生成し、長さ、解像度、フレームレート、音声生成の有無を選べます。
FLUX 3 Draft Extend: 720p での高速低コストのクリップ継続ドラフト
Minimax H3 Reference to Video は、書面によるプロンプトと提供された参照メディアから 768p または 2K ビデオを生成します。これは、繰り返し登場するキャラクター、製品、ブランドの外観など、複数のショットにわたって被写体を認識できるようにする必要がある作業向けに構築されています。少なくとも 1 つの参照画像またはビデオを添付してシーンを説明すると、モデルがそれらの参照が画面上でどのように動作するかを解決します。
最大 9 枚の参照画像、最大 3 つの参照ビデオ、最大 3 つの参照音声ファイルを添付できます。 Minimax H3 Reference to Video では、画像がアイデンティティ、スタイル、構成をガイドし、ビデオが動き、ペース、カメラの動作をガイドします。音声を単独で送信することはできません。画像またはビデオ参照を添付する必要があります。
Image-to-Video では、クリップの文字通りの最初のフレームが固定されるため、冒頭の構成が固定されます。 Minimax H3 Reference to Video は、代わりにメディアをガイドとして扱い、被写体の一貫性を保ちながらフレーミングやカメラの動きをより自由にします。正確な開始フレームよりも、誰が、または何が登場するかを重視する場合に選択してください。
出力解像度は「768p」、「2k」、または「として選択可能で、持続時間は 4–15 秒の間で選択できます。 Minimax H3 Reference to Video は、21:9、16:9、4:3、1:1、3:4、9:16 の 6 つのアスペクト比をサポートしているため、ワイドスクリーン、正方形、および垂直の成果物はすべて同じモデルから提供されます。現在公開されている値については、このページのパラメータ パネルを確認してください。
プロンプトは 1 ~ 7000 文字を受け入れ、リクエストを実行するには少なくとも 1 つの参照画像またはビデオが存在する必要があります。参照カウントの上限は、画像 9 個、ビデオ 3 個、オーディオ ファイル 3 個です。制限はモードまたはプロバイダーの設定によって異なる場合があるため、制限を回避するように構築する前にライブ パネルに対して確認してください。
参照間の関係を明確に述べます。どちらが主題で、どちらが動きを提供するかです。 Minimax H3 のビデオへの参照は、暗黙的な関係よりもはるかによく名前付きの関係に従い、競合する参照は互いに平均化される傾向があります。抽象的な雰囲気の言葉ではなく、目に見えるアクションやカメラの動作を説明します。
はい。参照、プロンプト、解像度、および継続時間が適切になるまで RunComfy AI Playground Web UI でプロトタイプを作成し、同じパラメーターを使用して RunComfy API を介して同じモデルを呼び出します。これにより、Minimax H3 Reference to Video は、独自のパイプライン内の手動探索と自動ジョブの間で同じ動作を維持します。
世代ごとに、768p の場合はカウント 1 秒あたり 0.10 ドル、2K は 1 秒あたり 0.16 ドルは 1 秒あたり。カウントされる秒数は、完成したクリップに添付した参照ビデオの合計時間を加えたものであるため、10 秒の 768p 出力と 5 秒の参照クリップの請求額は 15 秒、つまり 1.50 ドルになります。リファレンス クリップは実行後に測定されるため、送信前に表示される数値は推定値です。請求に関する質問については、hi@runcomfy.com までお問い合わせください。
RunComfyは最高の ComfyUI プラットフォームです。次のものを提供しています: ComfyUIオンライン 環境とサービス、および ComfyUIワークフロー 魅力的なビジュアルが特徴です。 RunComfyはまた提供します AI Models, アーティストが最新のAIツールを活用して素晴らしいアートを作成できるようにする。





