Downloads · 30 days
21
28% of all-time downloads
solvrays/scribegene-llm-v0.4
scribegene-llm-v0.4 is a image-to-text model from solvrays. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.
Vision-native handwritten insurance form understanding, fine-tuned from unsloth/Qwen2-VL-7B-Instruct using QLoRA.
Downloads · 30 days
21
28% of all-time downloads
All-time downloads
74
Public
Parameters
8.3B
16.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
Vision-native handwritten insurance form understanding, fine-tuned from unsloth/Qwen2-VL-7B-Instruct using QLoRA.
No OCR needed. This model reads handwriting, checks checkbox states, and extracts structured data directly from scanned MDF (Monthly Disability Verification) form images.
| Property | Value |
|---|---|
| Base Model | unsloth/Qwen2-VL-7B-Instruct (7B) |
| Task | Visual Question Answering on MDF forms |
| Fine-tuning Method | QLoRA (r=16, alpha=32) via Unsloth |
| Quantization | 4-bit NF4 (training) → 16-bit merged |
| Annotator | Vertex AI Gemini 2.5 Flash |
| Exact Match | 0% |
| OOD Refusal Rate | 0% |
| License | Apache 2.0 |
from transformers import AutoModelForCausalLM, AutoProcessor
from PIL import Image
import torch
model_id = "solvrays/scribegene-llm-v0.4"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="cuda",
trust_remote_code=True,
)
# Load your scanned MDF form image
image = Image.open("mdf_form.png").convert("RGB")
# Ask a question about the form
question = "What is the name of the physician who signed this form?"
messages = [{
"role": "user",
"content": [
{"type": "image"},
{"type": "text", "text": question}
]
}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to("cuda")
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=200, temperature=0.1)
answer = processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(answer)
A Monthly Disability Verification Form (Form 441.O.MDF.O) is issued by TriPlus Services, acting as Third-Party Administrator of Penn Treaty Network America and American Network policies. It requires a licensed physician to certify a patient's ongoing disability status monthly.
| Challenge | OCR Approach | This Model |
|---|---|---|
| Cursive physician names | Fails ("Carnazzo", "Kruszka") | Reads directly from image |
| Checkbox state (YES/NO) | Misses (no text to extract) | Sees the ✓/✗ mark in context |
| Date grid cells (MM/DD/YYYY) | Digit confusion in small boxes | Layout-aware reading |
| Signature field | Garbage output | Correctly ignored |
| Handwritten addresses | High error rate | Contextual correction |
Scanned MDF Form (PDF)
↓ Image pre-processing (deskew 300 DPI, bilateral denoise, CLAHE)
↓ Vertex AI Gemini 2.5 Flash → structured JSON annotation
↓ VQA triplet dataset (field extraction + OOD refusal pairs)
↓ Qwen2-VL-7B + QLoRA (Unsloth, 2-5× faster, 80% less VRAM)
↓ Merge adapters → full 16-bit model
↓ HuggingFace Hub (safetensors)
base_model: unsloth/Qwen2-VL-7B-Instruct
fine_tuning_method: QLoRA (NF4, double quantization)
lora_rank: 16
lora_alpha: 32
lora_dropout: 0.05
use_rslora: true
vision_layers: frozen
language_layers: adapted
optimizer: AdamW 8-bit (paged)
lr_scheduler: cosine
neftune_noise_alpha: 5
annotator: Vertex AI Gemini 2.5 Flash
framework: Unsloth + HuggingFace TRL
| Metric | Value |
|---|---|
| Exact Match (field extraction) | 0% |
| OOD Refusal Rate | 0% |
| Evaluation Set | Held-out MDF form pages |
OOD Refusal Rate measures how reliably the model declines to answer questions not answerable from the form (e.g. "What is the diagnosis?", "Has this claim been approved?").
null for blacked-out fields (insured name/policy number).This model is released under the Apache 2.0 License. The base model (unsloth/Qwen2-VL-7B-Instruct) is also Apache 2.0.