Downloads · 30 days
18
1% of all-time downloads
lucky-verma/driver-license-reader
driver-license-reader is a image-text-to-text model from lucky-verma. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
Downloads · 30 days
18
1% of all-time downloads
All-time downloads
2.7K
Public
Parameters
202M
810 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors809 MB · 99%
How the weights are stored.
F32202M · 100%
From the Hugging Face model README
Donut-based image-to-JSON extraction for driver's license documents.
driver license OCR alternative | ID card parsing | KYC document extraction | visual document understanding
</div>This model is a fine-tuned Donut VisionEncoderDecoderModel for extracting structured fields from driver's license images without a separate OCR pipeline. It converts an input image into JSON-like fields such as:
namestatedatedobpersonThe model is intended for demos, prototyping, and research around document AI workflows. It is not a production identity-verification system.
model.safetensors weight file and does not require loading Pickle weights.import re
import torch
from PIL import Image
from transformers import DonutProcessor, VisionEncoderDecoderModel
model_id = "lucky-verma/driver-license-reader"
processor = DonutProcessor.from_pretrained(model_id)
model = VisionEncoderDecoderModel.from_pretrained(model_id)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device).eval()
image = Image.open("redacted_or_synthetic_license.jpg").convert("RGB")
pixel_values = processor(image, return_tensors="pt").pixel_values.to(device)
task_prompt = "<s_cord-v2>"
decoder_input_ids = processor.tokenizer(
task_prompt,
add_special_tokens=False,
return_tensors="pt",
)["input_ids"].to(device)
with torch.inference_mode():
outputs = model.generate(
pixel_values,
decoder_input_ids=decoder_input_ids,
max_length=model.decoder.config.max_position_embeddings,
early_stopping=True,
pad_token_id=processor.tokenizer.pad_token_id,
eos_token_id=processor.tokenizer.eos_token_id,
use_cache=True,
num_beams=1,
bad_words_ids=[[processor.tokenizer.unk_token_id]],
return_dict_in_generate=True,
)
sequence = processor.batch_decode(outputs.sequences)[0]
sequence = sequence.replace(processor.tokenizer.eos_token, "")
sequence = sequence.replace(processor.tokenizer.pad_token, "")
sequence = re.sub(r"<.*?>", "", sequence, count=1).strip()
print(processor.token2json(sequence))
nielsr/donut-basemodel.safetensorsThis model was trained on a small driver's-license dataset and may fail on unseen layouts, glare, blur, occlusion, low-resolution scans, non-English documents, or non-US license formats. Treat the output as an extraction suggestion, not a verified identity record.