Downloads · 30 days
13
12% of all-time downloads
morriszjm/MiniMax-M2.5-tiny-24e
MiniMax-M2.5-tiny-24e is a text generation model from morriszjm. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Training-free expert-pruned variant of morriszjm/MiniMax-M2.5-tiny, produced by the minimaxexpertpruning pipeline.
Downloads · 30 days
13
12% of all-time downloads
All-time downloads
113
Public
Parameters
4.3B
5.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.5 GB · 100%
How the weights are stored.
F8_E4M33.1B · 71%
From the Hugging Face model README
Training-free expert-pruned variant of morriszjm/MiniMax-M2.5-tiny,
produced by the minimax_expert_pruning pipeline.
num_local_experts: 32 → 24 (pruning rate: 25.0 %)gate.weight and e_score_correction_bias per MoE layer are row-sliced to the
kept experts; per-expert tensors of dropped experts are absent; kept experts
are renumbered contiguously to 0..23.top_k = num_experts_per_tok is unchanged (8).We run a small calibration set (64 prompts spanning Nokia AI4Code, general English Q&A, multilingual, and reasoning) through the unpruned source model and hook every MoE layer's router. Per layer, we accumulate each expert's selected probability mass — the post-sigmoid routing weight that the expert receives, summed over all calibration tokens that selected it in their top-8. We keep the top-K by this score per layer (uniform K) and atomically slice the on-disk per-expert tensors. No gradients, no fine-tuning.
{"ai4code": 1008, "general_en": 416, "reasoning": 257, "multilingual": 170}Production-style serving of the source model's domain (Nokia / Merlin AI4Code plus general English) at reduced HBM footprint. Expect graceful quality degradation versus the unpruned source on tasks well-covered by the calibration mix; quality on out-of-distribution domains may drop further.
config.json, model-NNNNN-of-NNNNN.safetensors (FP8), model.safetensors.index.json,
tokenizer, custom modeling_minimax_m2.py + configuration_minimax_m2.py, and
expert_prune_plan.json (full record of which experts were kept per layer).
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("morriszjm/MiniMax-M2.5-tiny-24e", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"morriszjm/MiniMax-M2.5-tiny-24e", trust_remote_code=True,
torch_dtype=torch.bfloat16, device_map="auto",
)
For vLLM serving, pass --trust-remote-code and (on multi-GPU) match
--data-parallel-size to the EP topology you compiled the K against.