Downloads · 30 days
181
8% of all-time downloads
marksverdhei/MiniMax-M2.5-GGUF
MiniMax-M2.5-GGUF is a machine learning model from marksverdhei. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
GGUF quantizations of MiniMaxAI/MiniMax-M2.5, created with llama.cpp.
Downloads · 30 days
181
8% of all-time downloads
All-time downloads
2.2K
Public
Repo size
563 GB
Likes
3
Public
Click a slice to open those files.
.gguf563 GB · 100%
From the Hugging Face model README
GGUF quantizations of MiniMaxAI/MiniMax-M2.5, created with llama.cpp.
| Property | Value |
|---|---|
| Base model | MiniMaxAI/MiniMax-M2.5 |
| Architecture | Mixture of Experts (MoE) |
| Total parameters | 230B |
| Active parameters | 10B per token |
| Layers | 62 |
| Total experts | 256 |
| Active experts per token | 8 |
| Source precision | FP8 (float8_e4m3fn) |
| Quantization | Size | Description |
|---|---|---|
| Q8_0 | 227 GB | 8-bit quantization, highest quality |
| Q4_K_M | 129 GB | 4-bit K-quant (medium), good balance of quality and size |
| IQ3_S | 92 GB | 3-bit importance quantization (small), compact |
| Q2_K | 78 GB | 2-bit K-quant, smallest size |
These GGUFs can be used with llama.cpp and compatible frontends.
# Example with llama-cli
llama-cli -m MiniMax-M2.5-Q4_K_M.gguf -p "Hello" -n 128
float8_e4m3fn) precision, so Q8_0 is effectively lossless relative to the source weights.