MiniMax H3 Singularity Reference To Video: multi‑reference, identity‑preserving video generation in ComfyUI#
The MiniMax H3 Singularity Reference To Video ComfyUI workflow turns multiple character, style, and scene references into a coherent, temporally stable video. It is designed to preserve your subject across frames while allowing strong creative control over look, costume, lighting, and environment.
Built around a concise two‑stage diffusion recipe with a latent upscale in between, this workflow balances controlled motion with high fidelity detail. It is ideal for fashion films, stage performances, science‑fiction vignettes, and animated character sequences that must keep identity and styling consistent from start to finish.
Key models in Comfyui MiniMax H3 Singularity Reference To Video workflow#
- MiniMax‑h3 Singularity v1.3 (reference‑to‑video checkpoint) Community reference‑to‑video model used to fuse text prompts with image references and generate temporally coherent clips with strong identity and style preservation. Model card
- CLIP text encoder Maps prompts into a semantic space the video model understands, enabling prompt steering of composition, lighting, and motion. See the original paper for background on CLIP’s text‑image alignment: Learning Transferable Visual Models From Natural Language Supervision.
- Autoencoder (VAE) compatible with the checkpoint Encodes frames into latent space for fast diffusion and decodes final latents back to RGB while retaining fine detail.
How to use Comfyui MiniMax H3 Singularity Reference To Video workflow#
This ComfyUI graph follows a clear path: ingest references and prompts, generate a low‑cost base pass to lock identity and motion, upscale latents for detail, refine for sharpness, then assemble frames into a video.
References and prompt setup#
Provide multiple reference images that cover the role of your subject (front or 3/4 view for identity), a style frame that communicates art direction, and optional scene references for environment or palette. Keep faces unobstructed and reasonably sharp to help the model anchor identity. Write a concise positive prompt that names the subject, wardrobe, setting, and lighting, plus a short negative prompt to rule out artifacts you dislike. If you plan several shots, keep prompts stable and adjust only what must change to maintain continuity. The workflow blends all references with the prompt so the subject reads consistently across frames.
Base generation pass#
The base pass establishes subject likeness, pose, and broad motion while staying economical in steps. It trades micro‑detail for reliable identity and stable movement, which makes it well suited for iterative exploration. Set the seed for reproducibility across runs and tweak guidance strength only if the subject drifts. When you like the overall motion and framing, keep the seed fixed to carry the same identity into later stages. If the subject still wavers, strengthen the most relevant reference image or simplify the prompt.
Latent upscale#
The latent upscale stage enlarges the working resolution in latent space before fine detail is introduced. Upscaling at this point preserves the motion you approved in the base pass while creating headroom for crisp edges, hair detail, and fabric texture. Because the upscale happens on latents, it adds detail without introducing new motion jitter. If you intend to output higher resolutions, this stage is the lever that prepares clean pixels for refinement. Keep your seed and references unchanged to ensure continuity.
Refinement pass#
The refinement pass sharpens textures, contours, and small features without rewriting composition or motion. It is intentionally brief so it enhances rather than reimagines your base pass, which preserves identity and style from the references. Use this pass to tease out material properties like leather sheen, sequins, or sci‑fi paneling, and to reduce blur on faces and eyes. If you notice over‑sharpening or flicker, slightly lower prompt intensity or reduce how aggressively the refinement pass is applied. The output is a set of clean, consistent frames ready to be encoded.
Assembly and export#
The final stage bundles frames into a video clip at your chosen frame rate and resolution. Pick a duration and fps that suit your subject’s motion; slower, deliberate movement often looks best at moderate frame rates. Save to a common container such as MP4 or WebM for easy sharing. Name outputs methodically so you can compare variations by seed, references, or prompt notes. If you plan to edit shots together, export a few takes with the same seed to keep identity locked between cuts.
Optional extras#
- Curate references by role: 1–2 for identity, 1 for style, and 1 for environment. Avoid mixing conflicting lighting across references.
- Use tight crops for identity references so the face dominates the frame; use wider style or scene frames for look and palette.
- Keep prompts modular. Swap only the part that changes per shot, such as camera angle or background, to maintain identity.
- Start with short clips to validate motion and likeness, then scale duration once you are happy with the look.
- If you need stronger continuity between takes, reuse the same seed and keep the most influential reference image unchanged.
This MiniMax H3 Singularity Reference To Video workflow gives you a fast, repeatable way to turn curated references into cinematic, identity‑consistent video directly inside ComfyUI.
Acknowledgements#
This workflow implements and builds upon the following works and resources. We gratefully acknowledge RunningHub for the workflow source and WarmBloodAban for the MiniMax H3 Singularity model for their contributions and maintenance. For authoritative details, please refer to the original documentation and repositories linked below.
Resources#
- RunningHub/Workflow source
- Docs / Release Notes: Workflow source
- WarmBloodAban/Minimax-h3_Singularity
- Hugging Face: WarmBloodAban/Minimax-h3_Singularity
Note: Use of the referenced models, datasets, and code is subject to the respective licenses and terms provided by their authors and maintainers.


