Downloads · 30 days
8
15% of all-time downloads
mikrografija/doc-extractor-vl
doc-extractor-vl is a image-text-to-text model from mikrografija. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Document data extraction model based on Qwen2.5-VL-7B-Instruct, configured for structured JSON output from document images (invoices, forms, receipts, etc.).
Downloads · 30 days
8
15% of all-time downloads
All-time downloads
52
Public
Parameters
8.3B
16.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
Document data extraction model based on Qwen2.5-VL-7B-Instruct, configured for structured JSON output from document images (invoices, forms, receipts, etc.).
| File | Description |
|---|---|
cyrillic_logit_bias.json | 4129 token IDs with bias -100 to block Cyrillic generation |
system_prompt.txt | System prompt template for document extraction |
serving_config.yaml | Recommended vLLM serving parameters |
generate_cyrillic_bias.py | Script to regenerate the logit bias file |
vllm serve mikrografija/doc-extractor-vl --max-model-len 4096
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
# Load Cyrillic logit bias
with open("cyrillic_logit_bias.json") as f:
cyrillic_bias = {int(k): v for k, v in json.load(f).items()}
# Load system prompt
with open("system_prompt.txt") as f:
system_prompt = f.read()
response = client.chat.completions.create(
model="mikrografija/doc-extractor-vl",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."} },
{"type": "text", "text": "Extract data into this JSON schema: {\"issuer\": \"\", \"date\": \"\", \"total\": \"\", \"items\": []}"}
]}
],
logit_bias=cyrillic_bias,
temperature=0.0,
max_tokens=4096,
)
Qwen2.5-VL models are trained on multilingual data including Cyrillic scripts. When processing Latin-script documents (especially Slovenian, Croatian, or other languages with diacritics), the model occasionally substitutes Latin characters with visually similar Cyrillic characters (e.g., Latin "a" → Cyrillic "а"). The logit bias approach blocks this at the decoding level, making it impossible for the model to generate Cyrillic tokens.
This model uses unmodified Qwen2.5-VL-7B-Instruct weights. No fine-tuning was applied. The configuration files provide the Cyrillic blocking and structured output enforcement.