Divyarth Infotech
ComfyUIAugust 19, 2026Divyarth Infotech

MiniMax H3: Reference-to-Video with Audio in ComfyUI

1. Introduction

This tutorial focuses on the MiniMax H3 Reference-to-Video feature within ComfyUI, which lets you generate high-quality videos by providing one or more reference inputs — images, audio, or video — and letting the model build a new scene that stays consistent with what those references establish. In this guide we'll be working specifically with image references. The workflow is set up with three reference image loader nodes, but the examples in this tutorial use two reference images each. Keep in mind the underlying R2V model also accepts audio and video references for more advanced workflows.

With up to 2K output, a maximum duration of 15 seconds, and native stereo audio, MiniMax H3 Reference-to-Video gives creators a way to combine multiple source elements — a character, a location, a prop, a style — into a single cohesive shot.

By the end of this guide, you'll know how to set up the workflow, configure it, prompt it correctly for reference-driven generation, and produce your own Reference-to-Video clips with MiniMax H3 in ComfyUI.

MiniMax-H3 Available in The Hub Create cinematic videos with reference control and native stereo audio.

2. Requirements & Setup for MiniMax H3 Reference-to-Video (ComfyUI)

Getting MiniMax H3 Reference-to-Video up and running in ComfyUI starts with a handful of setup steps. Taking care of these first means the model is properly supported and the workflow runs without issues.

We're testing MiniMax H3 in this tutorial on an NVIDIA RTX 6000 Pro. Other GPUs will work too, but keep in mind that available VRAM plays a big role in the resolution, video length, and overall speed you can achieve.

Requirement 1: ComfyUI

You'll need ComfyUI running either locally or on a cloud GPU.

Local (Windows): 👉 How to Install ComfyUI Locally on Windows

Cloud GPU (RunPod): 👉 How to Run ComfyUI on RunPod with Network Volume

Requirement 2: Update ComfyUI

MiniMax H3 requires a recent version of ComfyUI with native support for the model. For this tutorial, make sure you are running ComfyUI version 0.30.0 or newer.

Keeping ComfyUI updated is important to ensure that the required H3 nodes and model support are available.

Windows Portable Users: Navigate to: ...\ComfyUI_windows_portable\update

Double-click: update_comfyui.bat

RunPod / Linux Users:

1 cd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspace
1 cd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspace

Alternatively, you can update ComfyUI directly through the ComfyUI Manager. After updating, restart ComfyUI and verify that you are running version 0.30.0 or newer.

Requirement 3: Download the Required MiniMax H3 Models

Next, you'll need to download the MiniMax H3 model files required for the Reference-to-Video workflow.

Below is a table showing the required files, their download pages, and where they should be placed inside your ComfyUI installation:

File NameDownload PageFolder
minimax_h3_ref2va_pruned_fp8_scaled.safetensors🤗 HuggingFace..\ComfyUI\models\diffusion_models
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors🤗 HuggingFace..\ComfyUI\models\text_encoders
minimax_h3_video_vae_fp16.safetensors🤗 HuggingFace..\ComfyUI\models\vae
minimax_h3_audio_vae_fp32.safetensors🤗 HuggingFace..\ComfyUI\models\vae

Requirement 4: Verify Your Folder Structure

Once all of the model files have finished downloading, verify that your ComfyUI folder structure matches the following:

1📂 ComfyUI/
2├── 📂 models/
3│   ├── 📂 diffusion_models/
4│   │   └── minimax_h3_ref2va_pruned_fp8_scaled.safetensors
5│   ├── 📂 text_encoders/
6│   │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
7│   └── 📂 vae/
8│       ├── minimax_h3_video_vae_fp16.safetensors
9│       └── minimax_h3_audio_vae_fp32.safetensors
1📂 ComfyUI/
2├── 📂 models/
3│   ├── 📂 diffusion_models/
4│   │   └── minimax_h3_ref2va_pruned_fp8_scaled.safetensors
5│   ├── 📂 text_encoders/
6│   │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
7│   └── 📂 vae/
8│       ├── minimax_h3_video_vae_fp16.safetensors
9│       └── minimax_h3_audio_vae_fp32.safetensors

With ComfyUI updated to version 0.30.0 or newer, all required MiniMax H3 model files downloaded, and the folder structure verified, you're ready to load the Reference-to-Video workflow and start generating videos with MiniMax H3.

3. Downloading and Loading the MiniMax H3 Reference-to-Video Workflow

Now that ComfyUI has been updated and all the required MiniMax H3 model files are in place, it's time to download and load the Reference-to-Video workflow into ComfyUI. The workflow includes the necessary nodes and settings to generate a video from one or more reference images with MiniMax H3.

Step 1: Download the Workflow

First, download the MiniMax H3 Reference-to-Video workflow JSON file. This file contains the complete workflow configuration and will allow you to quickly set up the generation process without having to build the workflow manually.

👉 Download MiniMax H3 Reference-to-Video Workflow JSON

Step 2: Load the Workflow

Once you have downloaded the workflow JSON file, open ComfyUI.

To load the workflow, simply drag and drop the JSON file into the ComfyUI interface. ComfyUI will automatically import the workflow and display all of the nodes and connections, including up to three separate reference image loader nodes.

Uploaded image

Step 3: Verify the Models

After loading the workflow, check that all required MiniMax H3 models have been detected correctly.

Verify the following components:

  • Diffusion Model: minimax_h3_ref2va_pruned_fp8_scaled.safetensors
  • Text Encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
  • Video VAE: minimax_h3_video_vae_fp16.safetensors
  • Audio VAE: minimax_h3_audio_vae_fp32.safetensors

If all four models are loaded correctly without any missing-model errors, the workflow is ready to use.

You can now move on to configuring the workflow settings and preparing your reference images for Reference-to-Video generation with MiniMax H3.

Reference to Video Available in The Hub

Guide your video with up to 3 reference images.

4. Configuring the Workflow Settings

The workflow is simple by design — load your references, write a prompt, set your resolution, and run:

Load Reference Images — in the Reference Images group, use the Load Image nodes to upload your reference images. For this tutorial we're using a maximum of three, though the model itself supports more if you want to extend the workflow. These can be a character, a location or background, a prop, an outfit, or a visual style — any combination of elements you want the model to draw from and combine into the generated video.

💡 Note: You don't need to fill all three reference slots. A single reference image lets the model animate that one source directly, while two or three references let the model combine multiple elements — for example, a character reference plus a separate environment reference — into one scene. The model can accept more than three references if your workflow needs them; three is simply the maximum we're using for the setup and examples in this tutorial.

Prompt — in the Prompt & Model Loading group, enter your prompt directly into the Reference to Video (MiniMax H3) node. This is where the four MiniMax H3 models are also loaded. We'll cover exactly how to write an effective reference-based prompt in Section 6.

Megapixel — in the Use Image Size group, set the megapixel value on the Scale Image to Total Pixels node. This controls your output resolution.

Once these are set, queue the prompt and the Output group's Save Video node will save the finished clip.

That's it — load references, prompt, megapixel, run.

Megapixel Setting

For our generations, we used 0.9 megapixels at 9:16. This produced an output resolution of approximately 736 × 1280 pixels — essentially 720p in a vertical format.

💡 Tip: If you want faster generations for testing prompts, drop the megapixel value to 0.4 (roughly 480p). It renders noticeably quicker and is a good way to iterate on a prompt before committing to a higher-resolution, slower generation at 0.9 or above.

Video Duration

The duration setting controls how long the generated video will be. Shorter clips are useful for testing how the model is interpreting and combining your references, while longer durations give MiniMax H3 more time to develop the scene once you're happy with the setup. The workflow automatically calculates the required frame count from the duration, so you don't need to manually configure the number of frames.

Final Settings

For the generations in this tutorial, we'll use:

SettingValue
Reference Images720 × 1280 pixels (9:16)
Mega Pixels0.9
Output Resolution736 × 1280 (9:16, ~720p)
Duration5 seconds
SeedRandomized
PromptCustom prompt describing how the references should combine, plus camera movement and audio

The remaining workflow settings can be left at their default values.

Once these settings are configured, you're ready to generate your first Reference-to-Video clip with MiniMax H3.

5. Generating Reference-to-Video Examples with MiniMax H3

With the workflow configured, let's look at Reference-to-Video in action.

For each example below, the reference images used are shown directly within the video itself, alongside the prompt used to generate the clip. Since Reference-to-Video is about combining source elements rather than transitioning between a start and end frame, seeing the references next to the final result makes it easy to tell exactly what came from where. For a breakdown of how these prompts are structured, see Section 6.

Example 1 — Clothing Transfer

Prompt

subject_definitions:
<Subject 1> is the woman in <Picture 1> and is the character identity reference. Use her exact face, facial features, hair, skin tone, body proportions, and overall appearance.
<Picture 2> is the complete fashion look reference for <Subject 1>. It shows the individual pieces of one coordinated outfit. Transfer all visible items from <Picture 2> onto <Subject 1>: the cream halter-style fringed top with its tie straps and metallic waist details, brown leather lace-up shorts with fringe details, brown western cowboy hat with decorative metal band, gold earrings, gold necklace, and tall brown-and-gold patterned western boots. Preserve the exact colors, materials, textures, patterns, shapes, proportions, and styling of each item.


summary:

Generate a 5-second ultra-realistic high-end western fashion campaign video of <Subject 1> wearing the complete fashion look from <Picture 2>. She confidently models the coordinated western outfit in a clean white professional studio with stylish CUTs showcasing the complete look and its individual fashion details.


detailed_description:

<Subject 1> remains fully recognizable throughout the entire video. Preserve her exact face, hair, skin tone, body proportions, and overall identity.


Style <Subject 1> in the complete look from <Picture 2>. Reproduce the cream fringed top, brown leather lace-up shorts, western cowboy hat, gold jewelry, and tall patterned western boots exactly as shown. Maintain consistent styling, proportions, materials, colors, and placement of every item throughout the video.

<Subject 1> moves like a professional fashion model with confident western-inspired movements: subtle weight shifts, elegant turns, a slow confident walk, and natural movement of the fringe, leather, and accessories.


0–0.7s: Full-body fashion shot. <Subject 1> is already wearing the complete look and takes a slow confident step toward the camera, clearly showing the entire outfit.


0.7–1.3s: CUT to a close-up of the upper body. Showcase the cream top, fringe, tie details, metallic waist details, and visible gold jewelry.


1.3–1.9s: CUT to a close-up of the brown leather shorts. Showcase the leather texture, lace-up details, fringe, stitching, and fit.


1.9–2.5s: CUT to a dedicated close-up of the boots and lower legs. Clearly showcase the tall western boots, their brown-and-gold pattern, texture, heel, and decorative details.


2.5–3.1s: CUT to a close-up of <Subject 1>'s face and hair. Clearly showcase the brown cowboy hat and gold earrings while she gives a confident western-fashion expression.


3.1–4.0s: CUT to a three-quarter full-body shot. <Subject 1> slowly turns while wearing the complete look, allowing the outfit and accessories to be seen from another angle.


4.0–5.0s: CUT to a final elegant full-body shot. <Subject 1> faces the camera with a confident western-fashion pose and subtle natural movement.


camera:

High-end contemporary western fashion campaign cinematography.


Use frequent, intentional CUTs with clearly different compositions. Use full-body framing for the complete look, close-ups for the top, shorts, jewelry, hat, and boots, and three-quarter framing to showcase the complete styling.


Use smooth controlled camera movement, precise focus, realistic leather and fringe physics, natural accessory movement, natural skin texture, shallow depth of field, and premium studio lighting.


Maintain accurate visual continuity of <Subject 1> and the complete fashion look from <Picture 2> across every CUT.


environment:

The background remains a clean, seamless pure-white professional photography studio throughout the entire video.


Use soft, premium fashion lighting with even illumination, realistic contact shadows, subtle highlights, and consistent exposure across every CUT.


overall_soundscape:

Sophisticated modern western fashion-campaign music plays throughout. Smooth deep groove, subtle percussion, atmospheric textures, and a confident rhythmic pulse. Stylish, luxurious, bold, and cinematic.


non_diegetic_music:

n/a

Result - (Reference images are shown in the video)

Example 2 — Product Reference

Prompt

 subject_definitions:
<Subject 1> is the woman in <Picture 1> and is the character identity reference. Use her exact face, facial features, hair, skin tone, body proportions, and overall appearance.
<Picture 2> is the aloe vera reference. Use it as the exact visual source for the fresh aloe vera plant and its natural contents. The only skincare substance used throughout the video is fresh aloe vera directly from <Picture 2>.


summary:

Generate a 5-second ultra-realistic playful skincare video of <Subject 1> using only fresh aloe vera from <Picture 2> inside a stylish modern bathroom. Keep <Subject 1> visible in every shot with multiple intimate close-up CUTs. She begins by saying “Facial time,” cleanly peels open the aloe vera leaf horizontally to reveal the fresh aloe vera gel, applies it generously over her entire face, and ends with a playful reaction as fresh aloe vera liquid drips from her nose.


detailed_description:

<Subject 1> remains fully recognizable throughout the entire video. Preserve her exact face, hair, skin tone, body proportions, and overall identity.


The aloe vera comes directly from the fresh leaf in <Picture 2>. It should look clear, watery, translucent, colorless to very pale green, glossy, slippery, and naturally fluid. The fresh aloe vera remains visibly liquid when applied to her face, forming a thin wet glossy layer while her natural skin remains visible underneath.


0–0.7s: Close chest-up frontal shot. <Subject 1> holds the fresh aloe vera leaf near her chest, looks directly into the camera, smiles playfully, and clearly says: “Facial time.”


0.7–1.5s: CUT to a tight three-quarter upper-body shot. <Subject 1> holds the aloe vera leaf horizontally with both hands and makes one clean horizontal peel across the top section of the leaf, pulling the green outer skin apart in one smooth motion. The opened section clearly reveals the fresh translucent aloe vera gel underneath. The gel looks wet, clear, glossy, and naturally fluid, stretching slightly and beginning to drip.


1.5–2.0s: CUT to a close face-and-hands shot. She takes the freshly exposed aloe vera directly from the opened leaf with her fingers and brings it toward her face.


2.0–2.6s: CUT to an intimate frontal close-up. She spreads a generous amount of fresh aloe vera across both cheeks with both hands. The aloe is visibly watery, transparent, glossy, and fluid.


2.6–3.2s: CUT to a slightly angled face close-up. She thoroughly massages the fresh aloe vera across her forehead, cheeks, nose, chin, jawline, and around her mouth, covering her entire face with the wet aloe.


3.2–3.8s: CUT to another tight frontal close-up. She continues rubbing the fresh aloe vera all over her face with both hands. Her entire face is visibly coated in a thin, transparent, glossy layer of fresh aloe, with realistic wet streaks and small droplets.


3.8–5.0s: CUT to an extreme close-up of her face. Her entire face remains wet and glossy with fresh aloe vera. She looks upward and rolls her eyes with exaggerated playful satisfaction. A noticeable amount of thin, transparent aloe vera liquid slowly drips from her nostrils, with only a few tiny liquid bubbles appearing naturally around her nose. She makes a brief playful, breathy reaction and momentarily appears to catch her breath before relaxing.


camera:

Use multiple smooth intentional CUTs while keeping <Subject 1> visible in every shot. Keep the camera very close, primarily framing her face, shoulders, chest, hands, and upper body.


Use frontal, three-quarter, and slightly angled close-ups with subtle handheld movement, gentle camera drift, shallow depth of field, precise focus, realistic skin texture, and detailed liquid physics.


Make the clean horizontal peeling action and the intensive aloe application the main visual focus. Clearly show the single horizontal peel revealing the fresh translucent aloe vera gel inside.


environment:

A stylish realistic modern indoor bathroom with a contemporary vanity and sink, large mirror, light-colored surfaces, warm neutral materials, tasteful bathroom fixtures, and soft ambient lighting. The bathroom feels upscale, clean, comfortable, and believable as a real skincare setting. Include realistic mirror reflections and subtle bathroom depth.


overall_soundscape:

Natural bathroom ambience, subtle aloe leaf peeling sound, wet aloe movement, soft hand-to-skin sounds, gentle breathing, and clear synchronized dialogue from <Subject 1> saying “Facial time” at the beginning. Include subtle wet dripping sounds and her brief breathy reaction during the final moment.


non_diegetic_music:

Light playful contemporary beauty-campaign music with a soft rhythmic groove and a humorous accent during the final reaction.

Result - (Reference images are shown in the video)

Reference to Video Available in The Hub

Guide your video with up to 3 reference images.

6. BONUS: How to Prompt for Reference-to-Video in MiniMax H3

Prompting effectively is crucial for achieving the desired results with the MiniMax H3 Reference-to-Video feature. A well-structured prompt can significantly enhance the model's ability to interpret your references and generate a cohesive video. Instead of writing a single block of text, break your prompt into labeled fields to provide clarity.

The five key fields to include are:

  1. subject_definitions — Clearly define each reference image and what it represents. Label them as Subject 1, Subject 2, etc., and provide concise descriptions of their key features.
  2. summary — Offer a brief overview of the target video, including its duration and the relationship between the subjects. This sets the stage for the detailed instructions that follow.
  3. detailed_description — This is where you provide shot-by-shot direction, including camera work, positioning, and interactions between subjects. It's beneficial to break this down into sections for clarity.
  4. overall_soundscape — Describe any diegetic sounds present in the scene, such as ambient noise or dialogue.
  5. non_diegetic_music — Specify any background music that is not part of the scene. If there is none, indicate this with n/a.

subject_definitions — Clearly define each reference image and what it represents. Label them as Subject 1, Subject 2, etc., and provide concise descriptions of their key features.

summary — Offer a brief overview of the target video, including its duration and the relationship between the subjects. This sets the stage for the detailed instructions that follow.

detailed_description — This is where you provide shot-by-shot direction, including camera work, positioning, and interactions between subjects. It's beneficial to break this down into sections for clarity.

overall_soundscape — Describe any diegetic sounds present in the scene, such as ambient noise or dialogue.

non_diegetic_music — Specify any background music that is not part of the scene. If there is none, indicate this with n/a.

By structuring your prompt in this way, you provide the model with a clear roadmap, enhancing its ability to generate a video that aligns with your creative vision. Remember to reference subjects by their labels throughout the prompt to avoid ambiguity.

7. Conclusion

In conclusion, the MiniMax H3 Reference-to-Video feature in ComfyUI offers a powerful way to build cinematic video content out of multiple source elements — characters, environments, objects, and style — combined into a single cohesive shot. Throughout this tutorial, we explored the essential steps, from setting up the environment and downloading the necessary files to configuring the workflow, prompting effectively for reference-driven generation, and producing videos.

The capabilities of MiniMax H3, including open weights and local generation, allow for high-quality outputs with native stereo audio, making Reference-to-Video an invaluable tool for content creators who want precise control over which elements appear in a shot and how they're combined.

As you continue to experiment with this technology, remember that practice is key. The more you work with different combinations of reference images, prompts, and framing, the better your results will become. Whether you are creating videos for social media, marketing, or personal projects, MiniMax H3's Reference-to-Video workflow provides the tools you need to bring separate source elements together into a single cohesive scene.

Enjoyed this article? Share it with your network.