Downloads · 30 days
0
PocketAiHub/MiniMax-Music3-MLX
MiniMax-Music3-MLX is a text-to-audio model from PocketAiHub. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
Experimental native Apple Silicon MLX inference for MiniMax-Music3. This repository runs the complete autoregressive, flow-DiT, and DAV synthesis path locally on macOS without CUDA or ComfyUI.
Downloads · 30 days
0
Access
Public
Updated Aug 15, 2026
Repo size
11.9 GB
Likes
5
Public
Click a slice to open those files.
.safetensors11.9 GB · 100%
From the Hugging Face model README
Experimental native Apple Silicon MLX inference for MiniMax-Music3. This repository runs the complete autoregressive, flow-DiT, and DAV synthesis path locally on macOS without CUDA or ComfyUI.
This is an independent community port, not an official MiniMax release. PocketAI did not train, fine-tune, or quantize the model weights; the packaged weights are unchanged from the pinned Comfy-Org repack identified below. No endorsement is implied.
The following one-minute rock-and-roll song was generated locally by this repository at 30 flow steps and seed 20260815.
<audio controls src="https://huggingface.co/PocketAiHub/MiniMax-Music3-MLX/resolve/main/examples/rock-and-roll-60s.wav"></audio>
Download the WAV · Generation parameters and signal checks
The acceptance render was produced as a 44.1 kHz, 16-bit stereo WAV. A 60-second song at 30 steps takes several minutes; exact speed depends on the Mac and available memory.
hf download PocketAiHub/MiniMax-Music3-MLX \
--local-dir MiniMax-Music3-MLX
cd MiniMax-Music3-MLX
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Create lyrics.txt:
[Verse]
Neon on the dashboard, midnight in the street
Engine keeps on rumbling to a backbeat
[Chorus]
Turn it up, let the good times roll
Fire in the speakers, thunder in your soul
Then run:
python generate.py \
--prompt "High-energy rock and roll, gritty male vocal, crunchy guitars, boogie piano, live drums, punchy bass, 148 BPM" \
--lyrics-file lyrics.txt \
--seconds 60 \
--steps 30 \
--seed 20260815 \
--output song.wav
For an instrumental, pass --lyrics "[Instrumental]". Supported duration is 10–300 seconds and supported flow-step count is 1–30.
| Component | File | Format |
|---|---|---|
| Global + local autoregressive model | text_encoders/minimax_music3_text_encoder_pruned_int8_convrot.safetensors | INT8 tensorwise + ConvRot |
| Flow diffusion transformer | diffusion_models/minimax_music3_dit_int8_convrot.safetensors | INT8 tensorwise + ConvRot |
| DAV waveform decoder | vae/minimax_music3_dav.safetensors | FP32 |
| Native runtime | minimax_mlx_model.py | MLX |
| Standalone CLI | generate.py | Python |
The weights are unchanged copies of the pinned Comfy-Org MiniMax-Music-3 repack. Exact sizes and SHA-256 checksums are recorded in model_manifest.json.
The runtime mirrors the MiniMax-Music3 implementation in ComfyUI commit efd4e951, including:
Long DAV decodes are processed using overlap-cropped safe-size chunks. Direct multi-million-sample MLX Conv1d execution produced incorrect channel collapse during testing; the chunked path is bit-for-bit identical to direct decoding at safe tensor sizes. The runtime also rejects outputs exhibiting the diagnosed stereo-collapse signature.
The included one-minute example passed these signal checks:
| Check | Result |
|---|---|
| Duration | 59.9888 seconds |
| Channel RMS | 0.1478 / 0.1501 |
| Channel peak | 0.9740 / 0.9900 |
| Stereo correlation | 0.7086 |
| Collapsed one-second blocks | 0% |
| Clipped samples | 0 |
Run the lightweight tests with:
python -m unittest tests/test_minimax_mlx_model.py
Model weights, this derivative package, and use of generated outputs are subject to the included MiniMax-Music3 Community License, including its acceptable-use policy and commercial-use terms. MiniMax-Music3 builds on Qwen3-8B and software components described in the upstream license.
Please review the license before downloading, redistributing, or deploying this repository. Users are responsible for ensuring they have the necessary rights to prompts, lyrics, reference material, and generated content.