Downloads · 30 days
621
30% of all-time downloads
eulogik/TinyDoc-VLM-256M
TinyDoc-VLM-256M is a visual question answering model from eulogik. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Downloads · 30 days
621
30% of all-time downloads
All-time downloads
2.1K
Public
Parameters
290M
1.2 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors1.2 GB · 95%
From the Hugging Face model README
⚠️ RETIRED RESEARCH CHECKPOINT — no performance claims
Measured 0.0% on OCRBench (n=1,000, all 10 categories) in 2026-09. The model degenerates under every prompt tried, including the exact training format (
"Extract document information: <image>"), with or without repetition controls; the LoRA adapter is worse (literal degeneration). The checkpoint is kept for reproducibility and research only.Earlier marketing numbers for this model (DocVQA ~65%, OCRBench ~60%) were never measured and have been withdrawn by the project. Evidence & artifacts:
evaluation/phase0/results/ocrbench_256m_full.*anddocs/BENCHMARKS.mdin the GitHub repo.
The working product of this project is the local grounded-extraction SDK
(pip install tinydoc) — schema-valid JSON, evidence bounding boxes, confidence
scores — running on free ollama:qwen2.5vl:3b, no API key:
| System | SROIE field F1 (n=100, same scorer) |
|---|---|
| TinyDoc pipeline (qwen2.5vl:3b) | 0.870 |
| PP-OCR + heuristics (free competitor) | 0.376 |
| Tesseract + regex | 0.227 |
All numbers recompute from committed artifacts
(python3 evaluation/phase0/recompute_scores.py).
from tinydoc.pipeline import ReceiptPipeline
pipe = ReceiptPipeline("auto") # free Ollama engine
doc = pipe.extract("receipt.jpg")
print(doc.fields, doc.schema_valid, doc.confidence)
for fr in doc.field_results:
print(fr.name, fr.value, fr.evidence["quote"], fr.evidence["bbox"])
Experimental 256M-parameter document VLM (2026-06 run), legacy 384px pipeline:
Image (384×384) → SigLIP-B/16 (93M) → PixelShuffle compressor ×3 (→64 tokens)
→ SmolLM2-135M decoder (30 layers, GQA, 8192 ctx) → multi-task heads
Verified facts (not claims):
image_token_id mismatch was fixed in evaluation code; it does not
explain the failures.TinyDoc-VLM-768-checkpoints) was abandoned: vision
tower dead at init.from tinydoc_vlm import TinyDocVLMForConditionalGeneration, TinyDocVLMProcessor
model = TinyDocVLMForConditionalGeneration.from_pretrained("eulogik/TinyDoc-VLM-256M")
processor = TinyDocVLMProcessor()
| Resource | URL |
|---|---|
| GitHub (eval artifacts, SDK, decision record) | github.com/eulogik/TinyDoc-VLM |
| PyPI SDK | pypi.org/project/tinydoc |
| LoRA adapter (also retired) | eulogik/TinyDoc-VLM-LoRA |
| Benchmark write-up (measured only) | docs/BENCHMARKS.md |
Apache 2.0. Free for commercial use.
Built by eulogik