Downloads · 30 days
35
1% of all-time downloads
jjmcarrascosa/vit_receipts_classifier
vit_receipts_classifier is a image classification model from jjmcarrascosa. Use it when you need a label for an image. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of google/vit-base-patch16-224-in21k on the cord, rvl-cdip, visual-genome and an external receipt dataset to carry out Binary Classification (ticket vs noticket).
Downloads · 30 days
35
1% of all-time downloads
All-time downloads
4K
Public
Parameters
85.8M
4.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.pt2.7 GB · 53%
From the Hugging Face model README
This model is a fine-tuned version of google/vit-base-patch16-224-in21k on the cord, rvl-cdip, visual-genome and an external receipt dataset to carry out Binary Classification (ticket vs no_ticket).
Ticket here is used as a synonym to "receipt".
It achieves the following results on the evaluation set, which contain pictures from the above datasets in scanned, photography or mobile picture formats (color and grayscale):
This model is a Binary Classifier finetuned version of ViT, to predict if an input image is a picture / scan of receipts(s) o something else.
Use this model to classify your images into tickets or not tickers. WIth the tickets group, you can use Multimodal Information Extraction, as Visual Named Entity Recognition, to extract the ticket items, amounts, total, etc. Check the Cord dataset for more information.
This model used 2 datasets as positive class (ticket):
cordhttps://expressexpense.com/blog/free-receipt-images-ocr-machine-learning-dataset/For the negative class (no_ticket), the following datasets were used:
RVL-CDIPvisual-genomeDatasets were loaded with different distributions of data for positive and negative classes. Then, normalization and resizing is carried out to adapt it to ViT expected input.
Different runs were carried out changing the data distribution and the hyperparameters to maximize F1.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | F1 |
|---|---|---|---|---|
| 0.0026 | 0.28 | 500 | 0.0187 | 0.9982 |
| 0.0186 | 0.56 | 1000 | 0.0116 | 0.9991 |
| 0.0006 | 0.84 | 1500 | 0.0044 | 0.9997 |