Downloads · 30 days
0
modulsx/MiniMax-Music-3-Turbo-FP8
MiniMax-Music-3-Turbo-FP8 is a text-to-audio model from modulsx. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
⚠️ Notice (v1.0 Community Release - Work in Progress / Experimental Calibration): This is a v1.0 calibration release created to accelerate local generation on NVIDIA RTX GPUs. Further refinements and community feedbac…
Downloads · 30 days
0
Access
Public
Updated Aug 22, 2026
Repo size
11.8 GB
Likes
2
Public
Click a slice to open those files.
.safetensors11.8 GB · 100%
From the Hugging Face model README
⚠️ Notice (v1.0 Community Release - Work in Progress / Experimental Calibration):
This is a v1.0 calibration release created to accelerate local generation on NVIDIA RTX GPUs. Further refinements and community feedback are welcome as we calibrate audio hyperparameters.
This repository contains the ultra-optimized FP8 (torch.float8_e4m3fn) weights and the Turbo 8-Step Step-Distillation LoRA for MiniMax Music 3, tailored specifically for lightning-fast inference in ComfyUI.
| Configuration | Full 190s Song Generation Time (RTX 4090) | Speedup | VRAM Footprint |
|---|---|---|---|
| Baseline (INT8 non-optimized) | ~895s (~15 minutes) | 1.0x | ~16 GB |
| Text Encoder & DiT FP8 | ~315s (~5 minutes 15s) | 2.8x faster | ~11 GB |
| 🔥 FULL STACK: FP8 + Turbo LoRA 8-Step + SageAttention | ⚡ ~245s (~4 minutes 05s) | 🚀 ~3.7x Faster (-11 min!) | ~11.5 GB |
| File Name | Size | Type | Destination Folder in ComfyUI |
|---|---|---|---|
minimax_music3_text_encoder_fp8_e4m3fn.safetensors | 8.48 GB | Pure FP8 Text Encoder (Clean, unrotated BF16 base) | ComfyUI/models/text_encoders/ |
minimax_music3_dit_fp8_e4m3fn.safetensors | 2.46 GB | Pure FP8 DiT Flow-Matching Model | ComfyUI/models/diffusion_models/ |
minimax_music3_turbo_lora_8step.safetensors | 180 MB | 8-Step Consistency Distillation LoRA (Rank 64 / Alpha 64) | ComfyUI/models/loras/ |
[Load Diffusion Model] (minimax_music3_dit_fp8_e4m3fn.safetensors)
│
▼
[Load LoRA] (minimax_music3_turbo_lora_8step.safetensors | strength: 0.85 ⭐ IDEAL)
│
▼
[Patch Sage Attention KJ] (sage_attention: auto)
│
▼
[KSampler]
⭐ LoRA Strength =
0.85(Recommended Sweet Spot):
Setting the LoRA strength to0.85(rather than1.0) provides the absolute best balance: it accelerates the diffusion sampling to 8-10 steps while letting the base model inject 100% crystal-clear vocal formants, diction, and delicate high-frequency percussion.
MiniMax Music3 Text Encodeclip : Connect to Load Text Encoder with minimax_music3_text_encoder_fp8_e4m3fn.safetensors.cfg_scale : 1.5 (Recommended: keeps 100% voice clarity, instrument separation and lyrical articulation).max_duration : Adjust to your desired song length (e.g. 120.0 for 2 min, 190.0 for 3 min 10s).KSamplersteps : 8 (Ultra-fast) or 10 - 12 (Studio Master Quality)sampler_name : eulerscheduler : simplecfg : 1.7denoise : 1.0minimax_music3_text_encoder_pruned_bf16.safetensors directly to float8_e4m3fn. All 166 sensitive layers (audio decoder heads, RMSNorms, token embeddings) remain in high-precision BF16 to guarantee zero early-stop anomalies and pitch accuracy.Training scripts and conversion tools are open-sourced at:
👉 GitHub: Guillaume-127/Minimax-music-3-Turbo-8-steps
Created by Guillaume-127. Released for the open-source AI music community.