Downloads · 30 days
40
0% of all-time downloads
SamMikaelson/deepseek-ocr-qvlm-4bit
deepseek-ocr-qvlm-4bit is a image-to-text model from SamMikaelson. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.
This is a 4-bit quantized version of deepseek-ai/DeepSeek-OCR using QVLM (Quantized Vision Language Model) technique, saved in SafeTensors format for easy deployment.
Downloads · 30 days
40
0% of all-time downloads
All-time downloads
48K
Public
Parameters
1.9B
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 100%
How the weights are stored.
I81.5B · 78%
From the Hugging Face model README
This is a 4-bit quantized version of deepseek-ai/DeepSeek-OCR using QVLM (Quantized Vision Language Model) technique, saved in SafeTensors format for easy deployment.
| Metric | Value |
|---|---|
| Original Size | 6363.12 MB (6.21 GB) |
| Quantized Size | 2199.39 MB (2.15 GB) |
| Size Reduction | 4165.03 MB (65.46%) |
| Compression Ratio | 2.89x |
| Format | SafeTensors |
import torch
from transformers import AutoModel, AutoTokenizer
from safetensors.torch import load_file
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(
"SamMikaelson/deepseek-ocr-qvlm-4bit",
trust_remote_code=True
)
# Load quantized model (weights only)
quantized_state_dict = load_file("model.safetensors")
# Note: You'll need to implement dequantization logic for inference
# The quantization metadata is stored in the safetensors metadata
from safetensors.torch import load_file, safe_open
import json
# Load model with metadata
model_path = "model.safetensors"
# Read metadata
with safe_open(model_path, framework="pt", device="cpu") as f:
metadata = f.metadata()
quantization_metadata = json.loads(metadata.get("quantization_metadata", "{}"))
# Load state dict
state_dict = load_file(model_path)
# Implement dequantization here based on metadata
# See QVLM repository for full implementation
model.safetensors - Quantized weights in SafeTensors format (2199.39 MB)config.json - Model configuration with quantization settingsquantization_config.json - Detailed quantization configurationquantization_results.json - Compression statisticstokenizer.json - Tokenizer vocabularytokenizer_config.json - Tokenizer configurationThe quantized model achieves 2.9x compression while maintaining similar accuracy to the original model. The 4-bit quantization significantly reduces memory requirements, making it suitable for deployment on resource-constrained devices.
This model uses QVLM (Quantized Vision Language Models) which applies:
@article{deepseek-ocr,
title={DeepSeek-OCR: Optical Character Recognition Model},
author={DeepSeek-AI},
year={2024}
}
@article{qvlm,
title={QVLM: Quantized Vision Language Models},
author={Wang, Changyuan},
year={2024},
url={https://github.com/ChangyuanWang17/QVLM}
}
This model inherits the Apache 2.0 license from the base DeepSeek-OCR model.
Quantized on 2026-01-05 using QVLM 4-bit quantization