Downloads · 30 days
49
36% of all-time downloads
appautomaton/MiniMax-Music3-MLX
MiniMax-Music3-MLX is a text-to-audio model from appautomaton. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
[](https://pypi.org/project/mlx-minimax-music3/) [](https://github.com/appautomaton/mlx-minimax-music3)
Downloads · 30 days
49
36% of all-time downloads
All-time downloads
136
Public
Parameters
2.4B
28.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors28.5 GB · 100%
From the Hugging Face model README
Precision-preserving MLX-native layout conversion of
MiniMaxAI/MiniMax-Music3
for local inference on Apple silicon. It is designed for use with
mlx-minimax-music3,
the independent pure-MLX inference project and Python package that generates
complete stereo music from lyrics and a structured music caption without
PyTorch, CUDA, or a cloud API at inference time.
This is a format and tensor-layout conversion. It is not trained, fine-tuned, merged, or quantized, and it does not claim authorship of the underlying model. MiniMax developed and released MiniMax Music 3; App Automaton converted the published checkpoint for the independent MLX runtime.
| Component | Stored dtype | Size | Role |
|---|---|---|---|
| Global language model | BF16 | 15.99 GiB | Long-range structure and semantic music tokens |
| RVQ depth decoder | BF16 | 1.20 GiB | Seven residual acoustic codebooks |
| Condition encoder | FP32 | 0.09 GiB | Continuous hidden-state fusion |
| Flow transformer | FP32 | 9.06 GiB | Flow-matching acoustic synthesis |
| Vocoder | FP32 | 0.20 GiB | Stereo waveform decode |
| Tokenizer, scheduler, and metadata | — | 0.01 GiB | Prompting and checkpoint contract |
The complete checkpoint is 26.56 GiB (28.52 GB decimal). The repository contains only the dense profile. It does not contain selective-q8 or persistent FP16 derivatives.
The conversion is pinned to official source revision
fbdf52fbaaca799592917417eb05f1899f1255ec.
Its manifest.json records the source revision, mapping version, component file
sizes, tensor counts, dtypes, and SHA-256 digests.
The converter reads and writes SafeTensors directly through MLX. PyTorch is not part of conversion or runtime inference.
Install the current package from
mlx-minimax-music3 on PyPI:
uv add --prerelease=allow mlx-minimax-music3
Download this checkpoint into a local weight directory:
hf download appautomaton/MiniMax-Music3-MLX \
--local-dir weights/mlx-dense/MiniMax-Music3
Generate one minute of instrumental melodic techno:
from mlx_minimax_music3 import (
GenerationRequest,
Music3Pipeline,
instrumental_lyrics,
)
pipeline = Music3Pipeline("weights/mlx-dense/MiniMax-Music3")
result = pipeline.generate(
GenerationRequest(
caption=(
"Global Metadata: melodic techno, 128 BPM, A minor, nocturnal and "
"cinematic, gradually rising energy. Vocal Details: instrumental, "
"no vocals. Arrangement: deep rounded kick, warm sub-bass, crisp "
"hats, syncopated percussion, analog arpeggiator, evolving pads, "
"a glassy bell motif, controlled builds, and a spacious final drop."
),
lyrics=instrumental_lyrics(
"intro", "groove", "build", "drop", "breakdown", "outro"
),
audio_duration=60.0,
seed=7,
),
output="outputs/melodic-techno.wav",
)
print(result.metadata.stage_timings)
print(result.metadata.memory_reports)
audio_duration is a ceiling because the model may emit its end token earlier.
Set min_audio_duration when a minimum frame count is required. The default
checkpoint path keeps the official mixed precision: BF16 autoregressive models
and FP32 acoustic models.
The runtime loads one stage at a time. Autoregressive models are released before the flow transformer is loaded, and acoustic models are released before final waveform decoding. This bounds unified-memory residency and avoids retaining the entire checkpoint in memory at once.
The current runtime writes native 44.1 kHz stereo PCM16 WAV. The official serving profile resamples its output to 32 kHz; reference-output parity for that final profile remains in progress.
This is an alpha release. The dense checkpoint has passed:
On an Apple M5 Max with 128 GB unified memory, the three-minute validation run took 18 minutes 16 seconds, peaked at approximately 19.93 GiB of process memory, and did not increase swap usage. This is one machine-specific observation, not a portable performance guarantee.
Listening validation across more prompts and seeds, long-form quality parity, and the reference 32 kHz output profile are still in progress.
This checkpoint is intended for local research, development, and music generation with the MLX runtime on Apple silicon.
For the original architecture description, prompt guidance, examples, and model
limitations, read the
MiniMaxAI/MiniMax-Music3 model card.
The converted checkpoint remains governed by the included
MiniMax-Music3 Community License,
including its attribution, acceptable-use, safeguards, and commercial terms.
Review that license before downloading, redistributing, or deploying the model.
The mlx-minimax-music3 runtime code is separately licensed under MIT.
MiniMaxAI/MiniMax-Music3appautomaton/mlx-minimax-music3mlx-minimax-music3 on PyPI