MiniMax Music 3: Generate AI Music in ComfyUI (TTM)
1. Introduction
In the realm of music creation, MiniMax Music 3 stands out as a powerful tool for generating complete songs using AI. This tutorial will guide you through the process of utilizing MiniMax Music 3 within ComfyUI, an open-source platform that allows for flexible and user-friendly music generation. Unlike traditional music generation models that often produce short clips, MiniMax Music 3 is designed to create full-length songs, up to five minutes in duration. This capability is particularly beneficial for musicians, composers, and content creators looking to generate original music that maintains coherence and structure throughout the entire piece.
The model operates by taking two primary inputs: lyrics and a music description. The lyrics dictate the words sung in the song, while the music description provides essential details about the genre, mood, instrumentation, and overall sound. This combination allows users to have significant control over the final output, ensuring that the generated music aligns with their creative vision.
In this tutorial, we will cover everything from setting up your environment to generating your first complete song. By the end, you will have a solid understanding of how to leverage MiniMax Music 3 to create unique musical compositions that can be tailored to your specific needs.
2. Requirements & Setup for MiniMax Music 3 in ComfyUI
Before diving into the music generation process, itβs crucial to ensure that your system meets the necessary requirements for running MiniMax Music 3 in ComfyUI. This section will outline the essential components and steps needed to set up your environment effectively.
Requirement 1: ComfyUI
You'll need ComfyUI running either locally or on a cloud GPU.
Local (Windows):
π How to Install ComfyUI Locally on Windows
Cloud GPU (RunPod):
π How to Run ComfyUI on RunPod with Network Volume
Requirement 2: Update ComfyUI
MiniMax H3 requires a recent version of ComfyUI with native model support. For this tutorial, make sure you are running ComfyUI 0.30.0 or newer.
Keeping ComfyUI up to date ensures that the required H3 nodes and model support are available.
Windows Portable Users
Navigate to:
...\ComfyUI_windows_portable\update
Then double-click:
update_comfyui.bat
RunPod / Linux Users
Run:
cd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspacecd /workspace/ComfyUI && git pull origin master && pip install -r requirements.txt && cd /workspaceYou can also update ComfyUI directly through the ComfyUI Manager.
Once the update is complete, restart ComfyUI and confirm that your installed version is 0.30.0 or newer.
Requirement 3: Download MiniMax Music 3 Model Files
Next, you will need to download the required model files for MiniMax Music 3. Below is a table listing the essential files and their respective download sources:
| File | Download | ComfyUI Folder |
|---|---|---|
| minimax_music3_dit_fp16.safetensors | π€ HuggingFace | ..\ComfyUI\models\diffusion_models |
| minimax_music3_text_encoder_pruned_int8_convrot.safetensors | π€HuggingFace | ..\ComfyUI\models\text_encoders |
| minimax_music3_dav.safetensors | π€ HuggingFace | ..\ComfyUI\models\vae |
Requirement 4: Verify the Folder Structure
After downloading the necessary files, itβs important to verify that your ComfyUI directory structure is set up correctly. The expected structure should look something like this:
1π ComfyUI/
2βββ π models/
3β βββ π diffusion_models/
4β β βββ minimax_music3_dit_fp16.safetensors
5β βββ π text_encoders/
6β β βββ minimax_music3_text_encoder_pruned_int8_convrot.safetensors
7β βββ π vae/
8β βββ minimax_music3_dav.safetensors1π ComfyUI/
2βββ π models/
3β βββ π diffusion_models/
4β β βββ minimax_music3_dit_fp16.safetensors
5β βββ π text_encoders/
6β β βββ minimax_music3_text_encoder_pruned_int8_convrot.safetensors
7β βββ π vae/
8β βββ minimax_music3_dav.safetensorsEnsure that each MiniMax Music 3 component is placed in its designated folder. Once you have updated ComfyUI and installed the required models, you will be ready to load the workflow and start generating music.
3. Downloading and Loading the MiniMax Music 3 Workflow
Once you have all the necessary model files in place, the next step is to download and load the MiniMax Music 3 workflow into ComfyUI. This workflow is essential as it contains the nodes and connections required to generate music without the need to build everything from scratch.
Step 1: Download the Workflow
To get started, you will need to download the MiniMax Music 3 ComfyUI Workflow JSON. This file contains all the configurations and settings needed to facilitate music generation. You can find the download link here:
π Download the MiniMax Music 3 ComfyUI Workflow JSON.
Step 2: Load the Workflow
After downloading the workflow JSON, follow these steps to load it into ComfyUI:

- Open ComfyUI on your system.
- Locate the downloaded workflow JSON file.
- Drag and drop the JSON file into the ComfyUI interface.
- Wait for the workflow to load completely.
Open ComfyUI on your system.
Locate the downloaded workflow JSON file.
Drag and drop the JSON file into the ComfyUI interface.
Wait for the workflow to load completely.
Step 3: Verify the Models
After loading the workflow, check that all required MiniMax Music 3 models have been detected correctly.
Verify the following components:
- Diffusion Model: minimax_music3_dit_fp16.safetensors
- Text Encoder: minimax_music3_text_encoder_pruned_int8_convrot.safetensors
- VAE: minimax_music3_dav.safetensors
If all three models are loaded correctly without any missing-model errors, the workflow is ready to use.
You can now move on to configuring the workflow settings and preparing your prompts for MiniMax Music 3 music generation.
4. Configuring Music Descriptions and Lyrics
Configuring the Music Description and Lyrics
MiniMax Music 3 provides two complementary inputs for controlling the generated song: Music Description and Lyrics. The Music Description is the first input field in the workflow and defines how the song should sound, while the Lyrics field defines the words that are sung and the vocal structure.
Step 1: Configre the Music Description
For more precise control, MiniMax Music 3 supports Fine-Grained Music Control through a Structured Caption with three sections:
- Global Metadata: genre, subgenre, BPM, key, scale, emotional progression, listening scenario and production profile.
- Vocal Details: vocal gender, timbre, performance style, harmony, backing vocals and vocal effects.
- Arrangement: primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures and spatial effects.
For example, a Structured Caption for a funky track could look like this:
Global Metadata:
Genre: Afro House / Deep House / Tech House
Subgenre: Funky Afro House / Afro Tech / Synth House
BPM: 124β126
Key: D minor
Scale: Dorian
Emotional Progression: Mysterious β Sexy β Groovy β Energetic β Euphoric
Listening Scenario: Late-night dancefloor / club
Production Profile: Deep, funky, warm, punchy and groove-focused, with prominent electronic saxophone.
Vocal Details:
Female lead: Deep, warm, breathy and sensual. Short, catchy and hypnotic.
Arrangement:
Primary instruments: Funky bass, four-on-the-floor kick, modern synths and electronic saxophone.
Secondary instruments: Afro percussion, congas, shakers, pads, arpeggios and synth stabs.
Evolution: Start sparse, build bass and drums, introduce brighter synths and active saxophone, strip back during the breakdown, then rebuild into the final groove.
Groove: Tight, bouncy and funky House rhythm.
Bass: Deep, warm and syncopated with a repeating funky groove.
Percussion: Dry congas, shakers, hand drums and crisp claps.
Textures: Warm synths, dreamy pads, filtered arpeggios, funky stabs and subtle acid movement.
Spatial effects: Wide synths, sidechain pumping, short vocal delays, atmospheric reverb and spacious saxophone echoes.Global Metadata:
Genre: Afro House / Deep House / Tech House
Subgenre: Funky Afro House / Afro Tech / Synth House
BPM: 124β126
Key: D minor
Scale: Dorian
Emotional Progression: Mysterious β Sexy β Groovy β Energetic β Euphoric
Listening Scenario: Late-night dancefloor / club
Production Profile: Deep, funky, warm, punchy and groove-focused, with prominent electronic saxophone.
Vocal Details:
Female lead: Deep, warm, breathy and sensual. Short, catchy and hypnotic.
Arrangement:
Primary instruments: Funky bass, four-on-the-floor kick, modern synths and electronic saxophone.
Secondary instruments: Afro percussion, congas, shakers, pads, arpeggios and synth stabs.
Evolution: Start sparse, build bass and drums, introduce brighter synths and active saxophone, strip back during the breakdown, then rebuild into the final groove.
Groove: Tight, bouncy and funky House rhythm.
Bass: Deep, warm and syncopated with a repeating funky groove.
Percussion: Dry congas, shakers, hand drums and crisp claps.
Textures: Warm synths, dreamy pads, filtered arpeggios, funky stabs and subtle acid movement.
Spatial effects: Wide synths, sidechain pumping, short vocal delays, atmospheric reverb and spacious saxophone echoes.This structured approach gives the model clear information about the style, vocals, instruments, groove and progression of the track while keeping the description organized and easy to modify.
Step 2: Configure the Lyrics
The Lyrics field defines the words that are sung. Lyrics may also include explicit section tags that help MiniMax Music 3 understand the intended song structure.
Supported section tags include:
[Intro]
[Verse]
[Pre-Chorus]
[Chorus]
[Post-Chorus]
[Bridge]
[Instrumental]
[Solo]
[Outro]For example:
[Intro]
hmm..
Something's calling
Hmmm... hahahahaha...
[Verse]
This higher stream...
feels way too real... hahahaha
I hope you can feel it too... hmmmmm
hmmmmm
[Pre-Chorus]
Don't be so shy... oh Why?
Oh why?
Come a little closer to me...Hahahaha...
aaaaaah aaaaaah oooeehhh ...
[Chorus]
Is this real?
Or just an illusion?
Take me higher...
[Outro]
Ooh... hahahahaHow the Two Inputs Work Together
The two fields serve different purposes:
- Music Description = how the song should sound.
- Lyrics = what is being sung and how the vocal sections are structured.
Start with the Music Description to establish the musical identity, vocals, instrumentation and production style. Then use the Lyrics field to define the vocal content and song structure. Together, these inputs provide much finer control over the final generation.
5. MiniMax Music 3 Examples
With everything configured, we are ready to showcase our first MiniMax Music 3 example.
For this example, we use the Music Description and Lyrics configured in the previous section to generate a one-minute song, demonstrating how MiniMax Music 3 combines the specified vocals, instruments, groove and overall arrangement into a complete musical track.
Listen to the MiniMax Music 3 example below π
Explore More MiniMax Music 3 Examples
MiniMax Music 3 supports a wide range of genres, so don't be afraid to experiment.
For more inspiration, check out the MiniMax Music 3 Demo Page. It includes more examples, longer generations, and the Music Descriptions used to create them, which can be useful when building your own Structured Captions.
6. Conclusion
MiniMax Music 3 brings full-song AI music generation directly into ComfyUI, giving you control over both the musical direction and the vocal content. By combining a structured Music Description with carefully written Lyrics, you can guide the genre, vocals, instrumentation, arrangement, mood, and overall production of your track.
In this guide, we covered the complete workflowβfrom installing the required models and loading the ComfyUI workflow to configuring Fine-Grained Music Control and generating your first song.
The best results come from experimenting. Try different genres, vocal styles, tempos, instruments, and arrangements to discover what works best for your sound.
Enjoyed this article? Share it with your network.
