Downloads · 30 days
27
1% of all-time downloads
GadflyII/MiniMax-M2.1-NVFP4
MiniMax-M2.1-NVFP4 is a text generation model from GadflyII. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
NVFP4 quantized version of MiniMaxAI/MiniMax-M2.1 for efficient inference on NVIDIA Blackwell GPUs.
Downloads · 30 days
27
1% of all-time downloads
All-time downloads
5.4K
Public
Parameters
129B
131 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors131 GB · 100%
How the weights are stored.
U8114B · 88%
From the Hugging Face model README
NVFP4 quantized version of MiniMaxAI/MiniMax-M2.1 for efficient inference on NVIDIA Blackwell GPUs.
| Property | Value |
|---|---|
| Base Model | MiniMaxAI/MiniMax-M2.1 |
| Architecture | Mixture of Experts (MoE) |
| Total Parameters | 229B |
| Active Parameters | ~45B (8 of 256 experts) |
| Quantization | NVFP4 (e2m1 format) |
| Size | 131 GB |
compressed-tensors with nvfp4-pack-quantized formatfrom vllm import LLM, SamplingParams
llm = LLM(
model="GadflyII/MiniMax-M2.1-NVFP4",
tensor_parallel_size=2,
max_model_len=4096,
gpu_memory_utilization=0.90,
trust_remote_code=True,
)
sampling_params = SamplingParams(
temperature=0.7,
top_p=0.9,
max_tokens=1024,
)
outputs = llm.generate(["Your prompt here"], sampling_params)
print(outputs[0].outputs[0].text)
Tested on 2x RTX PRO 6000 Blackwell (96GB each):
| Prompt Tokens | Output Tokens | Throughput |
|---|---|---|
| ~100 | 100 | ~73 tok/s |
| ~1260 | 1000 | ~72 tok/s |
Same as base model - see MiniMaxAI/MiniMax-M2.1 for details.