logo
RunComfy
  • ComfyUI
  • TrainerNew
  • Models
  • API
  • Pricing
discord logo
MODELS
Explore
All Models
LIBRARY
Generations
MODEL APIS
API Docs
API Keys
ACCOUNT
Usage

MiniMax H3 Max Reference to video: Multimodal Video on Models and API | RunComfy

minimax/minimax-h3-max/reference-to-video

MiniMax H3 Max Reference to video turns image, video, and audio references plus a prompt into 5-15s 480p/768p clips with native audio.

Table of contents

1. Get started2. Authentication3. API referenceSubmit a requestMonitor request statusRetrieve request resultsCancel a request4. File inputsHosted file (URL)5. SchemaInput schemaOutput schema

1. Get started

Use RunComfy's API to run minimax/minimax-h3-max/reference-to-video. For accepted inputs and outputs, see the model's schema.

curl --request POST \
  --url https://model-api.runcomfy.net/v1/models/minimax/minimax-h3-max/reference-to-video \
  --header "Content-Type: application/json" \
  --header "Authorization: Bearer <token>" \
  --data '{
    "prompt": "Image 1 is the near man; Image 2 is the far man. Keep both faces, hair, and jackets identical to their reference images. Dusk in an empty asphalt parking lot under soft sodium lights. They stand a few meters apart. The near man says flatly: \"So… you actually tried it?\" A beat of wind. The far man answers with a slow half-smile, warm and certain: \"Yeah. I love it.\" The near man exhales and shakes his head once, almost smiling. Static camera, medium wide. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement, no text, no watermark.",
    "prompt_expansion_mode": "balanced",
    "reference_images": [
      "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-1.webp",
      "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-2.webp"
    ]
  }'

2. Authentication

Set the YOUR_API_TOKEN environment variable with your API key (manage keys in your Profile) and include it on every request as a Bearer token via the Authorization header: Authorization: Bearer $YOUR_API_TOKEN.

3. API reference

Submit a request

Submit an asynchronous generation job and immediately receive a request_id plus URLs to check status, fetch results, and cancel.

curl --request POST \
  --url https://model-api.runcomfy.net/v1/models/minimax/minimax-h3-max/reference-to-video \
  --header "Content-Type: application/json" \
  --header "Authorization: Bearer <token>" \
  --data '{
    "prompt": "Image 1 is the near man; Image 2 is the far man. Keep both faces, hair, and jackets identical to their reference images. Dusk in an empty asphalt parking lot under soft sodium lights. They stand a few meters apart. The near man says flatly: \"So… you actually tried it?\" A beat of wind. The far man answers with a slow half-smile, warm and certain: \"Yeah. I love it.\" The near man exhales and shakes his head once, almost smiling. Static camera, medium wide. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement, no text, no watermark.",
    "prompt_expansion_mode": "balanced",
    "reference_images": [
      "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-1.webp",
      "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-2.webp"
    ]
  }'

Monitor request status

Fetch the current state for a request_id ("in_queue", "in_progress", "completed", or "cancelled").

curl --request GET \
  --url https://model-api.runcomfy.net/v1/requests/{request_id}/status \
  --header "Authorization: Bearer <token>"

Retrieve request results

Retrieve the final outputs and metadata for the given request_id; if the job is not complete, the response returns the current state so you can continue polling.

curl --request GET \
  --url https://model-api.runcomfy.net/v1/requests/{request_id}/result \
  --header "Authorization: Bearer <token>"

Cancel a request

Cancel a queued job by request_id; in-progress jobs cannot be cancelled.

curl --request POST \
  --url https://model-api.runcomfy.net/v1/requests/{request_id}/cancel \
  --header "Authorization: Bearer <token>"

4. File inputs

Hosted file (URL)

Provide a publicly reachable HTTPS URL. Ensure the host allows server-side fetches (no login/cookies required) and isn't rate-limited or blocking bots. Recommended limits: images ≤ 50 MB (~4K), videos ≤ 100 MB (~2–5 min @ 720p). Prefer stable or pre-signed URLs for private assets.

5. Schema

Input schema

{
  "type": "object",
  "title": "Input schema",
  "required": [
    "prompt",
    "prompt_expansion_mode",
    "reference_images"
  ],
  "properties": {
    "prompt": {
      "title": "Prompt",
      "description": "Describe the scene, motion, camera, and audio. Refer to assets as Image 1, Image 2, Video 1, Audio 1, and say what each reference supplies.",
      "type": "string",
      "default": "Image 1 is the near man; Image 2 is the far man. Keep both faces, hair, and jackets identical to their reference images. Dusk in an empty asphalt parking lot under soft sodium lights. They stand a few meters apart. The near man says flatly: \"So… you actually tried it?\" A beat of wind. The far man answers with a slow half-smile, warm and certain: \"Yeah. I love it.\" The near man exhales and shakes his head once, almost smiling. Static camera, medium wide. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement, no text, no watermark."
    },
    "reference_images": {
      "title": "Reference Images",
      "description": "Subject or style reference image URLs, cited in the prompt as Image 1, Image 2, and so on. Combined with videos and audios, at most 12 reference files.",
      "type": "array",
      "default": [
        "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-1.webp",
        "https://playgrounds-storage-public.runcomfy.net/tools/7453/media-files/ref-2.webp"
      ],
      "items": {
        "type": "string",
        "format": "image_uri"
      },
      "maxItems": 12
    },
    "reference_videos": {
      "title": "Reference Videos",
      "description": "Motion reference clips (about 2-15s each; combined duration at most 15s), cited as Video 1, Video 2. Combined with images and audios, at most 12 reference files.",
      "type": "array",
      "items": {
        "type": "string",
        "format": "video_uri"
      },
      "maxItems": 12
    },
    "reference_audios": {
      "title": "Reference Audios",
      "description": "Optional voice or ambience clips (about 2-15s each; combined duration at most 15s). Cannot be the only reference; provide at least one image or video with them.",
      "type": "array",
      "items": {
        "type": "string",
        "format": "audio_uri"
      },
      "maxItems": 12
    },
    "aspect_ratio": {
      "title": "Aspect Ratio",
      "description": "Output aspect ratio. adaptive follows the reference framing when possible.",
      "type": "string",
      "enum": [
        "adaptive",
        "21:9",
        "16:9",
        "4:3",
        "1:1",
        "3:4",
        "9:16"
      ],
      "default": "adaptive"
    },
    "resolution": {
      "title": "Resolution",
      "description": "Native generation resolution. 480p is faster/cheaper; 768p is the sharper default.",
      "type": "string",
      "enum": [
        "480p",
        "768p"
      ],
      "default": "768p"
    },
    "duration": {
      "title": "Duration",
      "description": "Length of the generated video in seconds (5-15).",
      "type": "integer",
      "minimum": 5,
      "maximum": 15,
      "default": 5
    },
    "prompt_expansion_mode": {
      "title": "Prompt Expansion Mode",
      "description": "How much effort to spend rewriting the prompt before generation. balanced returns quickly; quality spends longer on a richer prompt.",
      "type": "string",
      "enum": [
        "balanced",
        "quality"
      ],
      "default": "balanced"
    },
    "seed": {
      "title": "Seed",
      "description": "Fixed seed for reproducible results. Use -1 for a random seed.",
      "type": "integer",
      "default": -1
    },
    "enable_safety_checker": {
      "title": "Enable Safety Checker",
      "description": "If set to true, the safety checker will be enabled.",
      "type": "boolean",
      "default": true
    }
  }
}

Output schema

{
  "output": {
    "type": "object",
    "properties": {
      "image": {
        "type": "string",
        "format": "uri",
        "description": "single image URL"
      },
      "video": {
        "type": "string",
        "format": "uri",
        "description": "single video URL"
      },
      "images": {
        "type": "array",
        "description": "multiple image URLs",
        "items": {
          "type": "string",
          "format": "uri"
        }
      },
      "videos": {
        "type": "array",
        "description": "multiple video URLs",
        "items": {
          "type": "string",
          "format": "uri"
        }
      }
    }
  }
}
Follow us
  • LinkedIn
  • Facebook
  • Instagram
  • Twitter
Support
  • Discord
  • Email
  • System Status
  • Affiliate
Video Models
  • Wan 3.0 Prime Reference To Video
  • Wan 3.0 Prime
  • Wan 3.0 Prime Text To Video
  • Wan 2.6 Flash
  • MiniMax H3 Open Image to Video
  • MiniMax H3 Open
  • View All Models →
Image Models
  • Seedream 5.0 Pro Image Edit
  • Qwen Image 3.0 Pro Edit
  • Flux 2 Flash Edit
  • Nano Banana Pro
  • seedream 4.0
  • GPT Image 2 Image Edit
  • View All Models →
Legal
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
RunComfy
Copyright 2026 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.