Downloads · 30 days
63
100% of all-time downloads
mlboydaisuke/OvisOCR2-LiteRT
OvisOCR2-LiteRT is a image-text-to-text model from mlboydaisuke. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for litert-lm. The card lists the license as apache-2.0.
ATH-MaaS/OvisOCR2 (0.85B, apache-2.0) converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime. Requires litert-lm ≥ 0.15.
Downloads · 30 days
63
100% of all-time downloads
All-time downloads
63
Public
Repo size
1.3 GB
Likes
1
Public
Click a slice to open those files.
.litertlm1.3 GB · 100%
From the Hugging Face model README
ATH-MaaS/OvisOCR2 (0.85B, apache-2.0) converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime. Requires litert-lm ≥ 0.15.
OvisOCR2 is a page-level document-parsing model built by post-training Qwen3.5-0.8B: given a document page image it emits a Markdown transcription in reading order — text as Markdown, tables as HTML <table>, formulas as LaTeX. This package wires the checkpoint's own 12-layer ViT to LiteRT-LM's fast_vlm contract at a static 512×512 input, on the same rail as our litert-community/Qwen3.5-0.8B VL build (the two checkpoints share every config byte; only the weights differ).
| File | Recipe | Size |
|---|---|---|
OvisOCR2_int8.litertlm | int8 decoder (dynamic on linears + embedding; convs and the delta rule stay float, fp32 activations declared) + fp16 vision encoder / int8 adapter, static 512×512, six-signature prefill ladder | 1.30 GB |
Transcription of a rendered table page (512×512), CPU backend, greedy — output shown verbatim:
Quarterly Sales Report 2025
This report summarizes unit sales for the first three quarters. Phones remain the
strongest product line across all quarters.
<table border="1"><tr><td>Product</td><td>Q1</td>...</table>
Table 1: Unit sales by product and quarter.
$B$/B difference, the position-contract cost).transformers: the upstream repo has no
generation_config.json, so HF's eos is only <|endoftext|> and fp32 generation runs
through the turn end and rambles. This bundle declares both <|endoftext|> and
<|im_end|> as stop tokens and stops cleanly; add 248046 to eos_token_id on the HF
side before comparing.The OCR post-training visibly erodes general chat ability: on a generic 8-question sanity
gate the model tends to transcribe the prompt instead of answering it. The fp32 original
on the stock HF stack fails the same questions the same way — echoing 17 + 25 back,
looping Tokyo, drifting into a numbered list on a French-vocabulary question — measured
question-by-question against this build. It is the checkpoint's behavior, not a conversion
artifact. Feed it document pages with the transcription prompt below.
litert-lm run ./OvisOCR2_int8.litertlm \
--attachment page_512.png \
--prompt "
Extract all readable content from the image in natural human reading order and output the result as a single Markdown document. For charts or images, represent them using an HTML image tag: <img src=\"images/bbox_{left}_{top}_{right}_{bottom}.jpg\" />, where left, top, right, bottom are bounding box coordinates scaled to [0, 1000). Format formulas as LaTeX. Format tables as HTML: <table>...</table>. Transcribe all other text as standard Markdown. Preserve the original text without translation or paraphrasing." \
--backend cpu --cache no --temperature 0
The prompt is the upstream reference prompt verbatim. The runtime resizes the attachment to 512×512; pre-rendering your page at 512×512 keeps the aspect ratio under your control.
litert-lm benchmark (litert-lm 0.16.0), Apple M4 Max, -p 256 -d 256 --runs 3 --cache no, quiet machine:
| Backend | Prefill (256) | Decode | TTFT |
|---|---|---|---|
| GPU | 2175 tok/s | 140.9 tok/s | 0.13 s |
| CPU | 659 tok/s | 48.9 tok/s | 0.43 s |
On device (iPhone 17 Pro, iOS 27, litert-lm v0.16.0 vendored build, cold start, single runs):
| Backend | Decode | TTFT | Peak memory |
|---|---|---|---|
| GPU (Metal) | 48.5 tok/s | 1.5 s | 3.8 GB |
| GPU (Metal), image leg | 64.7 tok/s | 1.5 s | 3.8 GB |
| CPU | 18.0 tok/s | 1.0 s | 0.9 GB |
The six-signature ladder loads Metal cleanly (no jetsam, no maxNumTokens override), the
same envelope as the Qwen3.5-0.8B VL build. On-device image legs are grounded: shown a
poster, it starts transcribing the poster's title; shown a photo without text, it emits its
picture-region <img> tag — exactly its document convention.
<|endoftext|> (248044) and <|im_end|> (248046); the upstream repo has
no generation_config.json, so both were taken from config.json + the tokenizer and
verified against the tokenizer at bundle build.Conversion scripts and a step-by-step reproduction:
hf-to-litertlm (qwen35vl_work/).