Downloads · 30 days
0
0% of all-time downloads
TheHouseOfTheDude/MiniMax-M2.5
MiniMax-M2.5 is a text generation model from TheHouseOfTheDude. Use it when you need the model to write or continue text. It is set up for vllm. The card lists the license as other.
This repository contains quantized inference builds of MiniMaxAI/MiniMax-M2.5 exported in the compressed-tensors layout for vLLM.
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
743
Public
Repo size
269 GB
Likes
0
Public
Click a slice to open those files.
.md5.6 KB · 79%
From the Hugging Face model README
This repository contains quantized inference builds of MiniMaxAI/MiniMax-M2.5 exported in the compressed-tensors layout for vLLM.
MiniMax-M2.5 is a large Mixture-of-Experts (MoE) model. The attached quant scripts calibrate all experts (not just router top-k) to produce more robust scales across the full mixture.
This repo publishes two quant variants:
The
mainbranch is typically a landing page. The runnable artifacts live under the AWQ-INT4 and NVFP4 branches.
Each variant branch includes:
*.safetensors) + model.safetensors.index.jsonconfig.json with compressed-tensors quant metadataExports are written with save_compressed=True so vLLM can load them as compressed-tensors.
Calibration is MoE-aware:
Why it matters: If only top-k experts are exercised, rare experts can receive poor scales and quantize badly—leading to instability when those experts trigger at inference time.
The scripts are designed to quantize only the MoE expert MLP weights, e.g.:
block_sparse_moe.experts.*.w1block_sparse_moe.experts.*.w2block_sparse_moe.experts.*.w3Everything else is excluded for stability (embeddings, attention, router/gate, norms, rotary, lm_head, etc.).
num_bits=4, symmetric)lm_head (kept higher precision)lm_headBoth scripts use a dataset recipe YAML/config that controls:
max_seq_lengthnum_samplesTokenization behavior
padding=Falsetruncation=Truemax_length=MAX_SEQUENCE_LENGTHadd_special_tokens=FalseThe exact dataset names/counts live in your recipe file; this README documents the pipeline and knobs.
If the base ships FP8 parameters, the scripts:
pip install -U vllm
vllm serve TheHouseOfTheDude/MiniMax-M2.5:AWQ-INT4 \
--quantization compressed-tensors \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--dtype bfloat16
pip install -U vllm
vllm serve TheHouseOfTheDude/MiniMax-M2.5:NVFP4 \
--quantization compressed-tensors \
--tensor-parallel-size 8 \
--enable-expert-parallel
Notes
--max-model-len, batch size, and GPU memory utilization accordingly.vllm serve at the variant directory (e.g., .../AWQ-INT4 or .../NVFP4).Quantization changes weight representation only. It does not modify tokenizer, chat template, or safety behavior. Apply your own safety policies/filters as appropriate.