Downloads · 30 days
62
13% of all-time downloads
albertosei/layoutlmv3-receipt-parser
layoutlmv3-receipt-parser is a token classification model from albertosei. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
A fine-tuned LayoutLMv3 model for extracting structured information from receipt images with 89.34% validation accuracy.
Downloads · 30 days
62
13% of all-time downloads
All-time downloads
495
Public
Parameters
126M
504 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors504 MB · 99%
From the Hugging Face model README
A fine-tuned LayoutLMv3 model for extracting structured information from receipt images with 89.34% validation accuracy.
albertosei/layoutlmv3-receipt-parser| Metric | Value |
|---|---|
| Final Validation Accuracy | 89.34% |
| Training Loss (Epoch 1) | 0.6824 |
| Training Loss (Epoch 2) | 0.3278 |
| Validation Accuracy (Epoch 1) | 83.49% |
| Number of Entity Labels | 51 |
The model recognizes 25 entity types in BIO format:
vendor_name - Store/business namevendor_address - Physical addressvendor_phone_number - Contact numberdate - Transaction datetime - Transaction timereceipt_id - Receipt number/identifiercurrency - Currency typetotal_amount - Final totalsubtotal_amount - Subtotal before taxtax_amount - Tax amountservice_charge_amount - Service feesdiscount_amount - Discounts appliedtip_amount - Tip/gratuitycash_paid_amount - Cash paymentchange_amount - Change returnedcredit_card_amount - Credit card paymente_money_amount - Electronic paymentpayment_method - Payment typeline_item_name - Product/service nameline_item_quantity - Quantity purchasedline_item_unit_price - Price per unitline_item_total_price - Line item totalline_item_discount_amount - Item-level discountline_item_vat_status - VAT informationother - Miscellaneous informationfrom transformers import AutoProcessor, AutoModelForTokenClassification
from PIL import Image
import torch
# Load model and processor
processor = AutoProcessor.from_pretrained("albertosei/layoutlmv3-receipt-parser", apply_ocr=False)
model = AutoModelForTokenClassification.from_pretrained("albertosei/layoutlmv3-receipt-parser")
# Prepare inputs (requires external OCR for text and bounding boxes)
image = Image.open("receipt.jpg").convert("RGB")
words = ["STORE", "NAME", "Date:", "2024-01-01", "Total:", "25.99"] # From OCR
boxes = [[0, 0, 100, 20], [100, 0, 200, 20], [0, 20, 50, 40],
[50, 20, 150, 40], [0, 40, 50, 60], [50, 40, 150, 60]] # From OCR
# Process and predict
encoding = processor(image, words, boxes=boxes, return_tensors="pt")
with torch.no_grad():
outputs = model(**encoding)
predictions = outputs.logits.argmax(-1).squeeze().tolist()
# Convert to labels
predicted_labels = [model.config.id2label[pred] for pred in predictions]