Downloads · 30 days
80
14% of all-time downloads
vanch007/MiniMax-Music3-MLX-8bit
MiniMax-Music3-MLX-8bit is a text-to-audio model from vanch007. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
Native Apple MLX checkpoint for MiniMaxAI/MiniMax-Music3, converted with selective affine 8-bit quantization and BF16 exceptions.
Downloads · 30 days
80
14% of all-time downloads
All-time downloads
577
Public
Repo size
14.2 GB
Likes
2
Public
Click a slice to open those files.
.safetensors14.2 GB · 100%
From the Hugging Face model README
Native Apple MLX checkpoint for MiniMaxAI/MiniMax-Music3, converted with selective affine 8-bit quantization and BF16 exceptions.
Runtime source and installation instructions: vanch007/mlx-minimax-music3.
python -m pip install "mlx-minimax-music3[server] @ git+https://github.com/vanch007/[email protected]"
mlx-minimax-music3 generate \
--model vanch007/MiniMax-Music3-MLX-8bit \
--prompt "Warm acoustic pop with intimate female vocals and fingerpicked guitar." \
--lyrics $'[verse]\nMorning light across the sea\n[chorus]\nStay here and sing this song with me' \
--duration 10 \
--seed 7 \
--steps 30 \
--output song.wav
Apple silicon and Metal are required. The production runtime does not use PyTorch.
Quantized modules:
BF16 modules and parameters:
Each component directory includes a conversion manifest with per-shard SHA-256 values, source-shard coverage, tensor names, and remote verification metadata.
| Item | Revision |
|---|---|
| Source model | MiniMaxAI/MiniMax-Music3@c2509fd6b60d1ae169cd1df27f78a53174ba17e8 |
| Diffusers reference | huggingface/diffusers@c6da9936e4bda83107943a16eb8682e9a37d8527 |
| MLX runtime | vanch007/[email protected] |
The full provenance record is in source_manifest.json.
The complete repository was downloaded from its fixed Hugging Face commit and passed strict shard, SHA-256, tensor-name, and finite-value checks.
Fixed-input component comparisons against the official BF16 PyTorch reference produced:
| Component | Cosine similarity | Relative RMSE |
|---|---|---|
| Language model | 0.999822 | 0.025001 |
| RVQ depth decoder | 0.999916 | 0.012956 |
| Condition encoder | 0.999990 | 0.004654 |
| Flow transformer | 0.998990 | 0.045071 |
| Vocoder | 0.999877 | 0.015715 |
A local Apple M3 Max release run generated 10 seconds of audio using 250 autoregressive frames, two overlapping chunks, and 30 flow steps. The result was a finite, non-silent, unclipped 44.1 kHz stereo PCM WAV lasting 9.996 seconds. Peak MLX memory was 23.34 GiB.
These checks establish implementation and numerical alignment within the recorded quantized tolerances. Human listening review is pending; no claim of perceptual quality parity with the source release is made.
This converted checkpoint remains subject to the MiniMax-Music3 Community License, including its attribution, acceptable-use, safeguard, and commercial terms.