Divyarth Infotech
ComfyUISeptember 24, 2026Divyarth Infotech

MiniMax H3 Turbo: Image-to-Video with Custom Audio in ComfyUI

1. Introduction to MiniMax H3 Turbo with Custom Audio

This tutorial shows how to use MiniMax H3 Turbo in ComfyUI with a dedicated custom node that makes Image-to-Video with custom audio straightforward. You simply drop in any audio track — whether it’s a full song for the character to sing along to, a sensual ASMR voice recording, or an instrumental beat for her to dance to — and the workflow generates a cinematic video clip synced to your chosen audio. The custom node handles the integration cleanly, giving you full creative control over the sound and motion.

MiniMax-H3 Turbo Available in The Hub

Create cinematic videos with custom audio.

2. Requirements & Setup for MiniMax H3 Turbo with Custom Audio (ComfyUI)

Getting MiniMax H3 Turbo with Custom Audio up and running in ComfyUI starts with a handful of setup steps. Taking care of these first means the model is properly supported and the workflow runs without issues.

We're testing MiniMax H3 Turbo on an NVIDIA RTX 6000 Pro. Other GPUs will work too, but keep in mind that available VRAM plays a big role in the resolution, video length, and overall performance.

Requirement 1: ComfyUI

You'll need ComfyUI running either locally or on a cloud GPU.

Local (Windows): 👉 How to Install ComfyUI Locally on Windows

Cloud GPU (RunPod): 👉 How to Run ComfyUI on RunPod with Network Volume

Requirement 2: Update ComfyUI

Keeping ComfyUI up to date ensures you have the latest MiniMax H3 Turbo node support and bug fixes.

Windows Portable Users: Navigate to: ...\ComfyUI_windows_portable\update Double-click: update_comfyui.bat

1 cd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspace
1 cd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspace

RunPod / Linux Users: Use the command above. Alternatively, update via ComfyUI Manager and verify the new nodes appear in the node search.

Requirement 3: Download the Required Models

Download the following files and place them in the correct folders:

File NameDownload PageFolder
minimax_h3_fl2va_int8_convrot.safetensors🤗 HuggingFace..\ComfyUI\models\diffusion_models
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors🤗 HuggingFace..\ComfyUI\models\text_encoders
minimax_h3_video_vae_int8_convrot.safetensors🤗 HuggingFace..\ComfyUI\models\vae
minimax_h3_audio_vae_fp32.safetensors🤗 HuggingFace..\ComfyUI\models\vae
minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors🤗 HuggingFace..\ComfyUI\models\loras

Requirement 4: Verify Your Folder Structure

Confirm your files are in the right locations before loading the workflow.

1📂 ComfyUI/
2├── 📂 models/
3│   ├── 📂 diffusion_models/
4│   │   └── minimax_h3_fl2va_int8_convrot.safetensors
5│   ├── 📂 text_encoders/
6│   │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
7│   ├── 📂 vae/
8│   │   ├── minimax_h3_video_vae_int8_convrot.safetensors
9│   │   └── minimax_h3_audio_vae_fp32.safetensors
10│   └── 📂 loras/
11│       └── minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors
1📂 ComfyUI/
2├── 📂 models/
3│   ├── 📂 diffusion_models/
4│   │   └── minimax_h3_fl2va_int8_convrot.safetensors
5│   ├── 📂 text_encoders/
6│   │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
7│   ├── 📂 vae/
8│   │   ├── minimax_h3_video_vae_int8_convrot.safetensors
9│   │   └── minimax_h3_audio_vae_fp32.safetensors
10│   └── 📂 loras/
11│       └── minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors

You're now ready to load the MiniMax H3 Turbo workflow.

⚠️ Note: This workflow requires an additional custom node. Navigate to your ComfyUI/custom_nodes folder and run:

1 git clone https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef 
1 git clone https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef 

3. Running MiniMax H3 Turbo Workflows in ComfyUI

Now that ComfyUI has been updated and the required models are in place, it's time to download and load the MiniMax H3 Turbo first-frame workflow.

Step 1: Download the Workflow

👉 Download MiniMax H3 Turbo Custom Audio I2V Workflow JSON

Step 2: Load the Workflow

Drag the JSON file into ComfyUI or use the Load button. The graph should appear with all nodes connected.

minimax h3 turbo image to video with custom audio comfyui workflow next diffusion

Step 3: Load Your First Frame and Audio

Upload your starting image in the First Frame group and your custom audio file in the Audio Input group.

Step 4: Configure Your Prompt

MiniMax prompting works best when it clearly references the input image. Use this structure:

1
2For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
3
4integrated_multimodal_description: [Shot 1] Live-action close-up, the woman shown in <Picture 1> remains in the same tight framing, face, shoulders and upper chest filling the frame. Her eyes stay half-lidded and locked on camera with an intense, inviting gaze. She sings along with the lyrics — mouth shaping every word clearly and expressively — while moving strictly and sensually on the beat. The camera stays continuous and highly dynamic, never cutting: it fluidly shifts between gentle pushes in and pulls out, all landing exactly on the beat, always keeping her face and upper body fully visible and sharp while she sings and moves for the full duration.
5
6overall_soundscape: n/a
7
8non_diegetic_music: n/a
1
2For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
3
4integrated_multimodal_description: [Shot 1] Live-action close-up, the woman shown in <Picture 1> remains in the same tight framing, face, shoulders and upper chest filling the frame. Her eyes stay half-lidded and locked on camera with an intense, inviting gaze. She sings along with the lyrics — mouth shaping every word clearly and expressively — while moving strictly and sensually on the beat. The camera stays continuous and highly dynamic, never cutting: it fluidly shifts between gentle pushes in and pulls out, all landing exactly on the beat, always keeping her face and upper body fully visible and sharp while she sings and moves for the full duration.
5
6overall_soundscape: n/a
7
8non_diegetic_music: n/a

You can always extend or modify the prompt by adding camera motion, sensual dance moves, facial expressions, or any other creative details to better match your vision.

Step 5: Configure Settings and Run

Set your desired duration and resolution in the main settings. If you want the audio to start at a specific second (for example, skipping the intro of a song), open the H3 Song Audio + Masked Video Context custom node and adjust the clip_start_seconds value. This lets you precisely control where the audio track begins, ensuring perfect sync between the motion, singing, and your chosen music or voice clip. Once everything is configured, click RUN to generate the video with custom audio.

MiniMax-H3 Turbo Available in The Hub

Create cinematic videos with custom audio.

4. Examples: MiniMax H3 Turbo Image to Video with Custom Audio

Below you can find a video with different characters singing and dancing to the custom audio.

This showcase demonstrates multiple characters performing to synced audio tracks, highlighting how the workflow handles singing, dancing, and expressive motion from a single first frame.

6. Conclusion

In conclusion, this ComfyUI MiniMax H3 Turbo tutorial has shown you exactly how to run the MiniMax H3 Turbo image to video workflow with custom audio. You now know how to set up the MiniMax H3 Turbo LoRA, the updated MiniMax H3 Turbo video VAE, and the required custom node so you can generate first-frame videos synced to any audio track. Whether you want to create sound-synced video, character singing videos, or dance videos with your own music, the MiniMax H3 Turbo custom audio workflow in ComfyUI gives you full control. By following the steps in this MiniMax H3 Turbo tutorial you have everything needed to run MiniMax H3 Turbo in ComfyUI and produce high-quality image to video with custom audio results.

Enjoyed this article? Share it with your network.