Downloads · 30 days
42
100% of all-time downloads
Konthee/dots-mocr-nvfp4
dots-mocr-nvfp4 is a image-text-to-text model from Konthee. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This repository is a NVIDIA Model Optimizer quantized derivative of:
Downloads · 30 days
42
100% of all-time downloads
All-time downloads
42
Public
Parameters
2.4B
4.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.2 GB · 100%
How the weights are stored.
BF161.7B · 73%
From the Hugging Face model README
This repository is a NVIDIA Model Optimizer quantized derivative of:
dots-studio/dots.mocr
The checkpoint is intended for NVIDIA GPU deployment with runtimes that understand the ModelOpt Unified Hugging Face quantization format, especially vLLM.
This is a post-training-quantized model. Validate OCR/layout accuracy on your own document set before production use.
| Component | Precision / behavior |
|---|---|
| Language-model weights | NVFP4 |
| Language-model activations | NVFP4 (W4A4) |
| Vision tower | High precision / unquantized |
| Other multimodal components | High precision / unquantized |
| KV cache | Unquantized |
| ModelOpt qformat | nvfp4 |
| Export quantization algorithm | NVFP4 |
| PTQ language attention implementation | eager |
| PTQ vision attention implementation | sdpa |
This variant targets maximum FP4 compression and compute opportunity. It quantizes both language-model weights and language-model activations to NVFP4, while preserving the vision/multimodal side at high precision. It is the more aggressive of the two builds and should be benchmarked carefully for OCR accuracy.
ModelOpt's plain PTQ path for VLMs is used here intentionally.
The language model is quantized while the vision encoder and non-language multimodal components are kept at high precision.
No vision-quantization recipe is used.
This W4A4 build uses post-training calibration. The configured calibration size is 256 samples. CALIB_WITH_IMAGES=0; when disabled, calibration uses the normal text calibration path.
Build settings:
dots-studio/dots.mocr2565121eagersdpaflash_attn import does not block model loading.The following versions were used to produce this checkpoint:
2.8.0+cu1284.57.60.36.20.0.1.dev1+g87f7d1432Source revisions:
dots.mocr commit:
23f3e5612fb8066d4034d5ecfc8f33a9243533eb
NVIDIA Model Optimizer commit:
87f7d1432f6dccffe67069c84b9a18877a35019d
A recent vLLM release can load ModelOpt NVFP4 and
W4A16_NVFP4 checkpoints using:
modelopt_fp4
If this Hugging Face repository is private:
export HF_TOKEN="hf_xxx"
Then run:
docker run --rm \
--gpus all \
--ipc=host \
--shm-size=16g \
-p 8000:8000 \
-e HF_TOKEN \
vllm/vllm-openai:latest \
--model Konthee/dots-mocr-nvfp4 \
--quantization modelopt_fp4 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.90 \
--chat-template-content-format string \
--served-model-name model \
--trust-remote-code
For a public repository, -e HF_TOKEN can be omitted.
For W4A4 NVFP4, vLLM auto-selects an available FP4 backend. On Blackwell GPUs it can use native FP4-capable kernels such as CUTLASS or FlashInfer when available. On platforms without a supported native FP4 GEMM, vLLM may fall back to a weight-only execution path.
The PTQ staging checkpoint uses sdpa for the
Transformers vision module so that flash-attn is not required during export.
Recent vLLM releases have a native DotsOCRForCausalLM implementation, so
this Transformers-only staging choice is not the vLLM attention backend.
It is generally better to let vLLM select the linear backend automatically first.
Do not force a backend unless you have benchmarked it on your specific GPU and vLLM version.
curl http://localhost:8000/v1/models
Expected API endpoint:
http://localhost:8000/v1
dots.mocr should be used with the prompts provided by the upstream project for best document parsing behavior.
Images can be supplied as an HTTP URL or a base64 data URL.
Example request body:
{
"model": "model",
"messages": [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,<BASE64_IMAGE>"
}
},
{
"type": "text",
"text": "<DOTS_MOCR_PROMPT>"
}
]
}
],
"temperature": 0,
"max_tokens": 4096
}
For document parsing, use the prompt definitions provided by the upstream dots.mocr project rather than replacing them with a generic OCR prompt.
Useful upstream locations include:
dots_mocr/utils/prompts.py
demo/demo_vllm.py
dots_mocr/model/inference.py
This checkpoint intentionally uses the following structure:
Image
|
v
Vision tower
high precision
|
v
Multimodal projection / integration
high precision
|
v
Language model
NVFP4 W4A4
|
v
Output tokens
This allows the vision side of dots.mocr to remain at higher precision while reducing the memory / compute cost of the language model.
Quantization can affect:
For production use, compare this checkpoint against the original BF16 model on a representative validation set.
Recommended comparison:
BF16 original
vs
W4A16 NVFP4
vs
W4A4 NVFP4
Useful metrics include:
Tested on 2026-09-23 against the original dots-studio/dots.mocr endpoint using the
latest Runpod concurrency sweep (20260923T112115437972Z). This checkpoint is the W4A4
variant: the language-model weights and activations use NVFP4.
| Per-endpoint concurrency | Successful pages | Throughput (pages/s) | vs. FP16 throughput | p95 latency | CER vs. FP16 | Agreement (1 − CER) |
|---|---|---|---|---|---|---|
| 16 | 320/320 (100.00%) | 0.965 | +13.0% | 23.490 s | 1.0900% | 98.9100% |
| 32 | 319/320 (99.69%) | 1.069 | +8.8% | 41.895 s | 1.6155% | 98.3845% |
| 64 | 319/320 (99.69%) | 1.049 | +19.4% | 94.074 s | 7.2202% | 92.7798% |
Across the three disjoint 320-page sets, this model completed 958/960 requests (99.79%). The character-weighted CER was 3.3017% (96.6983% agreement) over successful FP16 reference outputs.
assets/fixtures. The three concurrency levels used disjoint image sets to avoid
vLLM prefix-cache reuse; within a level, all three model endpoints received the same
encoded images.InternalServerError on the same sample
ID. At concurrency 64, this model had one InternalServerError; no retries were made.For this run, concurrency 32 delivered the highest observed throughput (1.069 pages/s). Concurrency 64 was slower than 32 and had a much longer p95 latency. Consider validating on a representative, human-labeled OCR set before selecting a production configuration.
Native NVFP4 acceleration is most relevant on NVIDIA hardware with native FP4 support, especially Blackwell-class GPUs.
Runtime behavior on other NVIDIA GPU generations depends on the kernels available in the installed vLLM version.
This repository is a quantized derivative of:
dots-studio/dots.mocr
The original model's license and usage terms continue to apply.
Review the upstream model repository before redistribution or production use.