Downloads · 30 days
31
14% of all-time downloads
SamMikaelson/deepseek-ocr-mbq-w4bit
deepseek-ocr-mbq-w4bit is a feature extraction model from SamMikaelson. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
This is a fully standalone quantized version of deepseek-ai/DeepSeek-OCR using MBQ (Mixed-precision post-training quantization).
Downloads · 30 days
31
14% of all-time downloads
All-time downloads
220
Public
Parameters
3.3B
3.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.5 GB · 100%
How the weights are stored.
I83.2B · 95%
From the Hugging Face model README
This is a fully standalone quantized version of deepseek-ai/DeepSeek-OCR using MBQ (Mixed-precision post-training quantization).
✨ No need to download the original model - all architecture files included!
| Metric | Value |
|---|---|
| Original Size | 6,672 MB (6.67 GB) |
| Quantized Size | 3,510 MB (3.51 GB) |
| Size Reduction | 3,162 MB (47.4%) |
| Compression Ratio | 1.90x |
pip install torch transformers safetensors accelerate pillow
import torch
from transformers import AutoTokenizer, AutoModel
# Device setup
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load model and tokenizer directly - all files included!
tokenizer = AutoTokenizer.from_pretrained(
"SamMikaelson/deepseek-ocr-mbq-w4bit",
trust_remote_code=True
)
model = AutoModel.from_pretrained(
"SamMikaelson/deepseek-ocr-mbq-w4bit",
trust_remote_code=True,
torch_dtype=torch.bfloat16
)
# Load the quantized weights using the helper
from load_mbq_model import load_mbq_model
state_dict = load_mbq_model("./") # Assumes files are in current directory
model.load_state_dict(state_dict)
model = model.to(device).eval()
print("✅ Model loaded successfully!")
import torch
from transformers import AutoTokenizer, AutoModel
from safetensors.torch import load_file
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(
"SamMikaelson/deepseek-ocr-mbq-w4bit",
trust_remote_code=True
)
# Load quantized weights
state_dict = load_file("model.safetensors")
# Separate weights and scales
weights = {}
scales = {}
for name, param in state_dict.items():
if '.scale' in name:
scales[name.replace('.scale', '')] = param
else:
weights[name] = param
# Dequantize weights
dequantized_state_dict = {}
for name, param in weights.items():
if name in scales:
scale = scales[name]
dequantized = (param.float() * scale).to(torch.bfloat16)
dequantized_state_dict[name] = dequantized
else:
dequantized_state_dict[name] = param
# Load model architecture (included in this repo!)
model = AutoModel.from_pretrained(
"SamMikaelson/deepseek-ocr-mbq-w4bit",
trust_remote_code=True,
torch_dtype=torch.bfloat16
)
# Load the quantized weights
model.load_state_dict(dequantized_state_dict)
model = model.to(device).eval()
print("✅ Model loaded successfully!")
✅ Standalone: All files included, no need to download original model
✅ Smaller Size: 47% reduction in model size
✅ Easy Loading: Simple AutoModel.from_pretrained() with trust_remote_code=True
✅ Compatible: Works with standard transformers library
✅ Preserved Quality: Mixed-precision maintains model performance
MBQ (Mixed-precision post-training quantization) intelligently allocates different bit-widths to layers based on their sensitivity:
If you use this quantized model, please cite:
@misc{deepseek-ocr-mbq,
author = {SamMikaelson},
title = {DeepSeek-OCR MBQ Quantized Model},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/SamMikaelson/deepseek-ocr-mbq-w4bit}}
}
Original model:
@misc{deepseek-ocr,
title={DeepSeek-OCR},
author={DeepSeek-AI},
year={2024},
howpublished={\url{https://huggingface.co/deepseek-ai/DeepSeek-OCR}}
}
MIT License (same as the base model)
If you encounter issues loading the model:
trust_remote_code=True is setpip install -r requirements.txtload_mbq_model.py helper scriptFor questions or issues, please open an issue on the model repository.