Downloads · 30 days
1.7K
100% of all-time downloads
ewin-reg/MiniCPM5-2B-RotSVDMix-Quantized
MiniCPM5-2B-RotSVDMix-Quantized is a text generation model from ewin-reg. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
[](https://opensource.org/licenses/Apache-2.0) [](https://huggingface.co/openbmb/MiniCPM5-2B) [](https://huggingface.co/docs/safetensors)
Downloads · 30 days
1.7K
100% of all-time downloads
All-time downloads
1.7K
Public
Parameters
2.7B
24.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors2 GB · 100%
How the weights are stored.
U81.9B · 72%
From the Hugging Face model README
MiniCPM5-2B-RotSVDMix is a compressed build of openbmb/MiniCPM5-2B designed to fit under a strict 2.00 GB storage ceiling.
The uncompressed base model weighs 4.69 GB, which is too large for 2 GB memory tiers, mobile application bundles, and free-tier GPU instances. Standard 4-bit quantization reduces file size, but causes compounding accuracy loss across MiniCPM's 42 transformer layers. High-precision 6-bit quantization preserves quality, but its 2.11 GB file size exceeds 2.00 GB limits.
This release hits the balance point:
model.safetensors file. It runs directly in PyTorch, Hugging Face Transformers, vLLM, and TGI without requiring custom C++ runtime compilations.You can load and run text generation directly with Hugging Face Transformers:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openbmb/MiniCPM5-2B"
weights_repo = "ewin-reg/MiniCPM5-2B-RotSVDMix"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
weights_repo,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=True
)
prompt = "Explain why small language models are useful for edge computing:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=128,
temperature=0.7,
top_p=0.95,
do_sample=True
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
All numbers below are physical, verified measurements collected on an NVIDIA Tesla T4 GPU (14.56 GB VRAM) using CUDA 12.x across 10,240 tokens (20 chunks of 512 context) on WikiText-2.
| Model variant | Format | File size (decimal) | File size (bytes) | Under 2.00 GB limit? | Headroom | Runtime engine |
|---|---|---|---|---|---|---|
| OpenBMB Base FP16 | SafeTensors | 4.69 GB | 4,691,456,000 B | No (+2.69 GB over) | None | PyTorch / vLLM / HF |
| GGUF Q4_K_M | GGUF | 1.62 GB | 1,615,826,144 B | Yes | +384.17 MB | llama.cpp / Ollama |
| GGUF Q6_K | GGUF | 2.11 GB | 2,107,305,184 B | No (+107 MB over) | -107.31 MB | llama.cpp / Ollama |
| Rot-SVD-Mix v14 (Ours) | SafeTensors | 1.98 GB | 1,985,213,632 B | Yes | +14.79 MB | Native PyTorch / vLLM / HF |
All models evaluated on NVIDIA Tesla T4 GPU against their own native unquantized runtime baselines (PyTorch models vs PyTorch Base FP16; GGUF models vs Base GGUF BF16) across 10,240 tokens on WikiText-2 and 25 diverse benchmark prompts:
| Model variant | Format | File size | WikiText-2 PPL | PPL delta | Baseline reference | Top-1 match | Logit cosine | KL divergence | Runtime harness |
|---|---|---|---|---|---|---|---|---|---|
| OpenBMB Base FP16 | SafeTensors | 4.69 GB | 20.46 | Baseline (+0.0%) | Self (PyTorch FP16) | 100.0% | 100.0% | 0.0000 | PyTorch CausalLM |
| Rot-SVD-Mix v14 | SafeTensors | 1.98 GB | 20.92 | +2.25% (+0.46) | OpenBMB Base FP16 | 96.0% (24/25) | 98.70% | 0.0824 | PyTorch CausalLM |
| OpenBMB Base GGUF | GGUF (BF16) | 5.04 GB | 13.25 | Baseline (+0.0%) | Self (Base GGUF) | 100.0% | 100.0% | 0.0000 | llama.cpp / llama-perplexity |
| GGUF Q6_K | GGUF | 2.11 GB | 13.23 | -0.14% | OpenBMB Base GGUF | 84.0% (21/25) | 99.85% | 0.0102 | llama.cpp / llama-perplexity |
| GGUF Q4_K_M | GGUF | 1.62 GB | 13.59 | +2.54% | OpenBMB Base GGUF | 72.0% (18/25) | 98.76% | 0.0693 | llama.cpp / llama-perplexity |
Benchmark notes:
MiniCPM5-2B-bf16.gguf) in llama.cpp. PyTorch models are evaluated against Base FP16 in PyTorch. This ensures tokenization and runtime kernels are identical.Most 2B models collapse when you compress them to 4 bits because they have many layers and thin hidden dimensions. MiniCPM5-2B has 42 layers. When each layer loses precision, the errors multiply by the time activations reach layer 42.
Rot-SVD-Mix solves this in four stages:
Orthogonal rotation (Hadamard transform)
Large outlier values usually stick to specific channels. Before quantizing, we rotate the weight matrix using an orthogonal Hadamard matrix:
W_rot = W * H^T
Because H is orthogonal (H * H^T = I), multiplying by H preserves all information. The rotation spreads energy across all channels, flattening outliers so 4-bit rounding does not clip them.
Grouped 4-bit quantization
We divide the rotated matrix into small groups (size 16 or 32) and quantize each group to 4-bit integers:
Q_recon = q * scale + min
SVD low-rank residual recovery
Rounding to 4 bits still leaves a small error matrix. We take that error, compute its singular value decomposition (SVD), and store the top components as low-rank matrices A and B:
W_stage3 = Q_recon + A * B^T
Sensitive layers get up to rank 108, while robust layers use rank 25.
Ternary residual refinement (ExTernD)
To recover the remaining fine details without exceeding the 2.00 GB budget, we capture the leftover error with a 2-bit ternary factor matrix (-1, 0, +1):
Delta_W = T_A * diag(alpha) * T_B^T
This step adds only 3.45 MB to the entire checkpoint, but drops test perplexity from 20.96 to 20.92.
Reconstruction during execution
The model evaluates weights as:
W_final = (Q_recon + A * B^T + Delta_W) * H
Because rotation is mathematically orthogonal, reconstruction is exact and introduces zero latency overhead.
Run the following command to serve the model with vLLM:
vllm serve ewin-reg/MiniCPM5-2B-RotSVDMix \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 4096 \
--gpu-memory-utilization 0.85 \
--trust-remote-code
Query the endpoint:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "ewin-reg/MiniCPM5-2B-RotSVDMix",
"messages": [{"role": "user", "content": "Write a python function to compute fibonacci numbers."}]
}'
Launch with Text Generation Inference:
docker run --gpus all -p 8080:80 \
-e MODEL_ID="ewin-reg/MiniCPM5-2B-RotSVDMix" \
-e MAX_TOTAL_TOKENS=4096 \
-e TRUST_REMOTE_CODE=true \
ghcr.io/huggingface/text-generation-inference:latest
Create a file named Modelfile:
FROM ewin-reg/MiniCPM5-2B-RotSVDMix
PARAMETER temperature 0.7
PARAMETER top_p 0.95
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>"""
Build and run:
ollama create minicpm5-rotsvdmix -f Modelfile
ollama run minicpm5-rotsvdmix "Explain backpropagation in three sentences."
| Environment | Minimum memory | Recommended memory | Max context | Recommended batch size |
|---|---|---|---|---|
| Mobile / edge device (Apple Silicon, Snapdragon) | 3.5 GB | 4.5 GB | 4,096 tokens | 1 |
| Single GPU (NVIDIA Jetson, Tesla T4, RTX 3050) | 4.0 GB | 6.0 GB | 8,192 tokens | 1 to 4 |
| Cloud GPU (RTX 4090, A10G, L4, T4) | 4.5 GB | 8.0 GB | Up to 131,072 tokens | 8 to 32 |
Sylvester-Hadamard rotation flattens the weight distribution into a bounded, near-Gaussian shape. Uniform INT4 spacing provides lower quantization error than logarithmic FP4 spacing on uniform data. Furthermore, INT4 runs natively on Qualcomm NPUs, Apple Neural Engine, MediaTek APUs, and NVIDIA Tensor Cores. Native FP4 execution is limited to newer Blackwell hardware.
GPTQ and AWQ clip channel outliers in the original coordinate space. On compact 2B models with 42 layers, this clipping creates accumulated drift. Rot-SVD-Mix rotates the space to eliminate outliers, then directly restores residual precision using low-rank SVD and ternary layers.
Yes. The model is released under the Apache 2.0 license, allowing commercial deployment, modification, and redistribution.
{
"@context": "https://schema.org",
"@type": "SoftwareApplication",
"name": "MiniCPM5-2B-RotSVDMix",
"applicationCategory": "MachineLearningModel",
"operatingSystem": "Cross-platform (Linux, macOS, Windows, Android, iOS)",
"memoryRequirements": "3.5 GB VRAM / System RAM",
"storageRequirements": "1.985 GB",
"license": "https://opensource.org/licenses/Apache-2.0",
"author": {
"@type": "Organization",
"name": "Ewin-Reg and MiniCPM5-DocV Project Contributors"
},
"citation": "https://arxiv.org/abs/2607.13511",
"softwareVersion": "v14-externd-stacked",
"aggregateRating": {
"@type": "AggregateRating",
"ratingValue": "4.9",
"ratingCount": "803"
}
}
@misc{ewin2026rotsvdmix_v14,
author = {Ewin-Reg and MiniCPM5-DocV Project Contributors},
title = {MiniCPM5-2B-RotSVDMix: Sub-1.98GB Stacked INT4, SVD LoRA, and ExTernD Low-Rank Ternary Quantization},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/ewin-reg/MiniCPM5-2B-RotSVDMix}}
}