Downloads · 30 days
10
42% of all-time downloads
rachitpandey26/obj_v1
obj_v1 is a image-text-to-text model from rachitpandey26. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Vision-language model fine-tuned for structured data extraction from Indian financial documents. Give it a page image and a JSON schema; it returns the schema filled in from what is on the page. A 4B vision-language m…
Downloads · 30 days
10
42% of all-time downloads
All-time downloads
24
Public
Parameters
4.4B
8.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README
Vision-language model fine-tuned for structured data extraction from Indian financial documents. Give it a page image and a JSON schema; it returns the schema filled in from what is on the page. A 4B vision-language model, LoRA fine-tuned and merged. Nothing extra is needed at load time -- it is a plain bf16 checkpoint.
vllm serve objectai/obj_v1 \
--served-model-name obj_v1 \
--max-model-len 16384 \
--limit-mm-per-prompt '{"image":1}' \
--mm-processor-kwargs '{"max_pixels":1003520}' \
--trust-remote-code
max_pixels is 1280x28x28, the resolution the model was trained at. Raising it
wastes KV cache; lowering it makes small print unreadable.
The server is OpenAI-compatible, so an ordinary chat completion works:
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
"date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
model="obj_v1",
temperature=0.0,
max_tokens=8192,
messages=[
{"role": "system", "content":
"You are a document data extraction model. "
"Extract only values present in the document. "
"Use null for fields that are absent or illegible. "
"Output a single compact JSON object matching the requested schema. "
"No prose, no markdown, no explanation."},
{"role": "user", "content": [
{"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text",
"text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
]},
],
)
print(response.choices[0].message.content)
Match training or accuracy drops. The system prompt above is verbatim, and the user turn is the image followed by exactly two lines:
document_type: <type>
schema: <compact json>
Set temperature=0.0 so the same page yields the same answer.
| Weights | 8.9 GB (bf16) |
| VRAM | 16 GB minimum, 24 GB comfortable |
| Precision | bf16 (Ampere or newer; use fp16 below that) |
| Context | 16384 covers the longest documents |
Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add --dtype float16. | |
Long documents matter: bank_statement and form16 answers run to ~2500 | |
tokens, so max_tokens below 4096 truncates them mid-JSON. |
Compact JSON matching the requested schema. Fields absent from the page come
back null rather than guessed. Values found on the page that the schema did
not ask for are placed under extras when that key is included in the schema.
Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.