Downloads · 30 days
42
2% of all-time downloads
sayed0am/Nanonets-OCR2-3B-FP8-Dynamic
Nanonets-OCR2-3B-FP8-Dynamic is a image-text-to-text model from sayed0am. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Downloads · 30 days
42
2% of all-time downloads
All-time downloads
1.7K
Public
Parameters
4.1B
5.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.4 GB · 100%
How the weights are stored.
F8_E4M32.8B · 68%
From the Hugging Face model README
Creation Code
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor.utils import dispatch_for_generation
MODEL_ID = "nanonets/Nanonets-OCR2-3B"
# Load model.
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(MODEL_ID, torch_dtype="auto")
processor = AutoProcessor.from_pretrained(MODEL_ID)
# Configure the quantization algorithm and scheme.
# In this case, we:
# * quantize the weights to fp8 with per channel via ptq
# * quantize the activations to fp8 with dynamic per token
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8_DYNAMIC",
ignore=["lm_head", "re:visual.*", "re:model.visual.*"],
)
# Apply quantization and save to disk in compressed-tensors format.
oneshot(model=model, recipe=recipe)