Downloads · 30 days
19
10% of all-time downloads
st0722/OvisOCR2-MLX-4bit
OvisOCR2-MLX-4bit is a image-text-to-text model from st0722. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
这是 ATH-MaaS/OvisOCR2 的非官方 MLX 4-bit 量化版本,面向 Apple Silicon Mac 上的 MLX-VLM 和 oMLX。
Downloads · 30 days
19
10% of all-time downloads
All-time downloads
182
Public
Parameters
853M
645 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors625 MB · 96%
How the weights are stored.
U32752M · 88%
From the Hugging Face model README
这是 ATH-MaaS/OvisOCR2 的非官方 MLX 4-bit 量化版本,面向 Apple Silicon Mac 上的 MLX-VLM 和 oMLX。
This is an unofficial MLX 4-bit quantized conversion of ATH-MaaS/OvisOCR2 for document OCR and document parsing on Apple Silicon Macs. It is a conversion and quantization of the upstream checkpoint, not a new training run or a fine-tuned checkpoint.
| Item | Value |
|---|---|
| Base model | ATH-MaaS/OvisOCR2 |
| Model family | Qwen3.5 VLM (model_type: qwen3_5) |
| Main use | OCR, document parsing, Markdown extraction |
| MLX format | Affine 4-bit |
| Quantization | bits=4, group_size=64, mode=affine |
| Processor | Qwen3VLProcessor |
| Weight size | Approximately 625 MB for model.safetensors |
| Conversion status | Community conversion; not affiliated with the upstream authors |
Use this checkpoint to extract readable content from document images, including:
The 4-bit checkpoint is intended to reduce memory and storage requirements for local Apple Silicon inference. Quantization can cause small quality differences from the BF16 version, especially on tiny characters, dense tables and difficult layouts. The output should be validated on the document types that matter to you.
The simplest runtime is mlx-vlm on an Apple Silicon Mac:
python3 -m pip install -U mlx-vlm huggingface_hub
For a clean environment, use a virtual environment instead of installing packages into the system Python:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip mlx-vlm huggingface_hub
MLX requires Apple Silicon. CUDA, ROCm and ordinary x86 CPU inference are not the target runtime for this repository.
Replace YOUR_HF_USERNAME with the account or organization that publishes this repository:
hf download YOUR_HF_USERNAME/OvisOCR2-MLX-4bit \
--local-dir ./OvisOCR2-MLX-4bit
The repository should contain the extracted MLX model files at its root. Do not put the model inside another nested directory, and do not upload only the .tar archive.
The official mlx-vlm command is mlx_vlm.generate:
mlx_vlm.generate \
--model ./OvisOCR2-MLX-4bit \
--image /path/to/document.png \
--prompt "Extract all readable content from this image in Markdown. Preserve the original reading order, headings, paragraphs, lists, and table structure as much as possible. Return Markdown only; do not add explanations." \
--max-tokens 4096 \
--temperature 0.0
For OCR, thinking is normally unnecessary. Do not pass --enable-thinking unless you intentionally want to test a thinking-style prompt.
A shorter prompt can also be used:
Extract all readable content from the image in Markdown. Preserve the original text and layout as much as possible. Return Markdown only.
oMLX treats this checkpoint as a Qwen3.5 vision-language model. Copy the model into the directory configured for oMLX, then start the server:
mkdir -p ~/models
hf download YOUR_HF_USERNAME/OvisOCR2-MLX-4bit \
--local-dir ~/models/OvisOCR2-MLX-4bit
omlx serve --model-dir ~/models
Open http://localhost:8000/admin/chat, select the model and upload a document image. If an oMLX installation does not identify the model automatically, set its model type to VLM in the Admin panel. Use a prompt like the one above and keep thinking disabled for normal OCR.
The first request can be slower because the model and Metal resources have to be loaded. Subsequent requests are the more useful measure of inference speed. Cold-start time also depends on whether oMLX has evicted the model, the current memory pressure, the storage device and the installed oMLX/MLX version.
The conversion was performed from the upstream Hugging Face checkpoint with mlx-vlm:
mlx_vlm.convert \
--hf-path ATH-MaaS/OvisOCR2 \
--mlx-path OvisOCR2-MLX-4bit \
--quantize \
--q-bits 4
The resulting MLX configuration declares affine 4-bit weights with group size 64:
{
"quantization": {
"group_size": 64,
"bits": 4,
"mode": "affine"
}
}
Re-running the conversion with a different mlx-vlm version may produce small metadata differences. Keep the generated config.json, processor files and tokenizer files together with the quantized weights. The command-line options of newer mlx-vlm releases should be checked with mlx_vlm.convert --help.
The model repository contains the MLX quantized weight file and the configuration, tokenizer and processor files required by mlx-vlm/oMLX. Typical files include:
README.md
config.json
model.safetensors
processor_config.json
preprocessor_config.json
tokenizer.json
tokenizer_config.json
chat_template.jinja
The exact auxiliary filenames may vary slightly with the mlx-vlm release. Do not rename or remove files generated by the converter.
Observed warm-run results on one Apple Silicon Mac were approximately 141–155 generated tokens/second, but this is not a universal benchmark. Startup time and throughput depend on the Mac model, image token count, output length, memory pressure, cache state and software versions. The model card intentionally makes no general speed guarantee.
The upstream model page identifies ATH-MaaS/OvisOCR2 as Apache-2.0. Please review and comply with the upstream license and attribution requirements when redistributing this conversion. This repository is an unofficial conversion and quantization and does not change the upstream license.
Upstream resources:
This repository contains the 4-bit MLX conversion named OvisOCR2-MLX-4bit. It should be used together with the model files in this repository, not with the original Transformers/PyTorch weights directly.