Downloads · 30 days
0
kengboon/sam3_openvino
sam3_openvino is a machine learning model from kengboon. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for openvino. The card lists the license as other.
Pre-exported OpenVINO IR variants of SAM3.1 for efficient CPU and GPU inference.
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
12.7 GB
Likes
0
Public
Click a slice to open those files.
.bin10.9 GB · 75%
From the Hugging Face model README
Pre-exported OpenVINO IR variants of SAM3.1 for efficient CPU and GPU inference.
These models are derived from Meta's SAM 3.1 (facebook/sam3), and therefore follow the same license under the SAM License
The OpenVINO IR exports and quantized variants in this repository are derivative works of the original SAM 3.1 model weights and are subject to the same SAM License terms. No model architecture was modified — only the format (PyTorch → ONNX → OpenVINO IR) and optional weight compression were applied.
| Variant | Description | Size |
|---|---|---|
openvino-fp16 | FP16 — recommended for GPU | 1.63 GB |
openvino-fp32 | FP32 — reference precision for CPU | 3.25 GB |
openvino-int8_sym | INT8 symmetric — CPU-optimized | ~0.83 GB |
openvino-int8_asym | INT8 asymmetric — CPU-optimized | ~0.83 GB |
openvino-int4_sym | INT4 symmetric — ultra-low memory CPU | ~0.45 GB |
openvino-int4_asym | INT4 asymmetric — ultra-low memory CPU | ~0.45 GB |
openvino-int8_w8a16 | W8A16 weight-only — recommended for both CPU and GPU | 0.84 GB |
openvino-int8_ptq_gpu | W8A8 post-training quantization — best CPU throughput | ~0.83 GB |
| Variant | Potatoes | Candies | Nuts | LED | Mean | vs FP16 |
|---|---|---|---|---|---|---|
| FP16 | 473 | 450 | 488 | 398 | 452 | 1.00× |
| W8A16 | 495 | 474 | 512 | 419 | 475 | 0.96× |
| PTQ W8A8 | ~645 | ~620 | ~670 | ~550 | ~621 | 0.73× |
| Variant | Potatoes | Candies | Nuts | LED | Mean | vs FP16 |
|---|---|---|---|---|---|---|
| FP16 | 638 | 752 | 449 | 535 | 593 | 1.00× |
| W8A16 | 657 | 773 | 472 | 556 | 614 | 0.96× |
| Variant | Potatoes | Candies | Mean | vs FP16 | F1 Potatoes | F1 Candies |
|---|---|---|---|---|---|---|
| FP32 | 8,525 | 8,407 | 8,466 | 1.03× | 1.000 | 0.992 |
| FP16 | 8,862 | 8,648 | 8,755 | 1.00× | 1.000 | 0.992 |
| INT8 sym | 5,270 | 5,170 | 5,220 | 1.67× | 1.000 | 0.992 ✓ |
| INT8 asym | 5,360 | 5,213 | 5,286 | 1.66× | 1.000 | 0.992 ✓ |
| W8A16 | 5,228 | 5,183 | 5,205 | 1.68× | 1.000 | 0.992 ✓ |
| INT4 sym | 5,526 | 5,450 | 5,488 | 1.59× | 0.963 | 0.792 ⚠️ |
| INT4 asym | 5,599 | 5,485 | 5,542 | 1.58× | 1.000 | 0.939 ⚠️ |
| Dataset | FP16 (text) | W8A16 (text) | FP16 (canvas) | W8A16 (canvas) |
|---|---|---|---|---|
| Potatoes | 1.000 | 1.000 | 1.000 | 1.000 |
| Candies | 0.992 | 0.992 | 0.991 | 0.991 |
| Nuts | 0.698 | 0.703 | 0.874 | 0.881 |
| LED | 0.000* | 0.000* | 0.720 | 0.724 |
* LED dataset uses visual-only prompts; F1=0.000 in text mode is expected (no text labels provided).
| Approach | Avg text (ms) | Avg canvas (ms) | vs FP16 | Notes |
|---|---|---|---|---|
| FP16 (baseline) | 452 | 593 | 1.00× | Native GPU precision |
| PTQ W8A8 (all layers) | ~640 | ~780 | 0.73× | Q/DQ overhead on FP16-optimized GPU |
| PTQ W8A8 (VE only) | ~650 | ~760 | 0.72× | Mixed-precision boundary overhead |
| W8A16 (weight-only) | 475 | 614 | 0.96× | No activation Q/DQ; near-lossless |
Why PTQ is slower on GPU: Modern GPU compute units (e.g., FP16 DPAS) are highly optimized for FP16 math. NNCF PTQ inserts explicit Q/DQ activation nodes at every layer boundary; those extra kernel launches exceed the memory savings from 2× smaller weights. W8A16 (weight-only) has no such nodes — weights are dequantized on-the-fly in a fused kernel — giving only ~4% overhead.
Why INT8 is faster on CPU: VNNI instructions natively accelerate INT8 dot products with no separate Q/DQ overhead. The OV CPU plugin fuses dequantize into the matmul kernel, giving 1.67× speedup with zero accuracy loss.
| Target hardware | Recommended variant | Reason |
|---|---|---|
| GPU (XPU / discrete) | openvino-fp16 | Fastest; native FP16 compute |
| GPU (memory-constrained) | openvino-int8_w8a16 | 2× smaller, only 4% slower |
| CPU | openvino-int8_w8a16 | 1.68× faster than FP16, lossless |
| CPU (ultra-low memory) | openvino-int4_sym | Smallest, but F1 may drop on some datasets |
| Older iGPU | openvino-int8_sym | Better for architectures without optimized FP16 |
nncf.compress_weights, PTQ: nncf.quantize)from instantlearn.models.sam3 import SAM3OpenVINO
from instantlearn.models.sam3.sam3_openvino import SAM3OVVariant
# FP16 — fastest on GPU
model = SAM3OpenVINO(variant=SAM3OVVariant.FP16, device="GPU")
# W8A16 — recommended for CPU or memory-constrained GPU
model = SAM3OpenVINO(variant=SAM3OVVariant.INT8_W8A16, device="CPU")
model = SAM3OpenVINO(variant=SAM3OVVariant.INT8_W8A16, device="GPU")