Downloads Β· 30 days
0
ishank9/llava_id_extraction
llava_id_extraction is a image classification model from ishank9. Use it when you need a label for an image. It is set up for pytorch. The card lists the license as apache-2.0.
Document Classification Module of the LLaVA ID Extraction Project
Downloads Β· 30 days
0
Access
Public
Updated Jul 25, 2026
Repo size
346 MB
Likes
1
Public
Click a slice to open those files.
.pth346 MB Β· 100%
From the Hugging Face model README
Document Classification Module of the LLaVA ID Extraction Project
This repository contains a fine-tuned Vision Transformer (ViT-Base) model that simultaneously predicts:
using a shared Vision Transformer backbone with three independent classification heads.
This model serves as the first component of the larger LLaVA ID Extraction project, whose goal is to build an end-to-end AI pipeline for automated identity document understanding, information extraction, and verification.
Current Pipeline
Identity Document
β
βΌ
Multi-Head Vision Transformer β
β
βββ Document Type
βββ Front / Back
βββ State
β
βΌ
OCR (Planned)
β
βΌ
LLaVA Information Extraction (Planned)
β
βΌ
Identity Verification (Planned)
β
βΌ
Fraud Detection (Planned)
Shared Vision Transformer encoder with three task-specific classification heads:
The model was trained on a synthetic dataset containing Indian identity document images with annotations for:
The dataset includes multiple document categories and image variations suitable for supervised multi-task learning.
| Parameter | Value |
|---|---|
| GPU | NVIDIA Tesla T4 |
| Epochs | 4 |
| Optimizer | AdamW |
| Scheduler | Cosine Annealing |
| Mixed Precision | FP16 |
| Image Size | 224 Γ 224 |
| Backbone | ViT Base Patch16 224 |
| Task | Accuracy |
|---|---|
| Document Type | 100% |
| Document Side | 100% |
| State | 100% |
| Combined Prediction | 100% |
These results were obtained on the held-out validation split of the synthetic dataset.
multihead_vit_best.pth
label_mappings.json
training.ipynb
import torch
model.load_state_dict(
torch.load(
"multihead_vit_best.pth",
map_location="cpu"
)
)
model.eval()
This model represents the document classification module of the broader LLaVA ID Extraction project.
Future additions include:
The model was trained exclusively on synthetic identity document images.
Although it achieves excellent performance on the held-out synthetic validation set, additional fine-tuning and evaluation on real-world document scans and photographs are recommended before production deployment.
GitHub repository: