Downloads · 30 days
4
3% of all-time downloads
mkdigitalgmbh/runpo-LayoutLM3-Invoice-Receipt
runpo-LayoutLM3-Invoice-Receipt is a token classification model from mkdigitalgmbh. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
LayoutLMv3 model initialized for receipt and invoice field extraction.
Downloads · 30 days
4
3% of all-time downloads
All-time downloads
157
Public
Parameters
126M
504 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors504 MB · 99%
From the Hugging Face model README
LayoutLMv3 model initialized for receipt and invoice field extraction.
⚠️ This is an initialized base model - not yet fine-tuned on custom data.
microsoft/layoutlmv3-baseThis model is configured to extract the following fields from receipts and invoices:
[ "O", "B-MerchantName", "I-MerchantName", "B-MerchantAddress", "I-MerchantAddress", "B-TransactionDate", "I-TransactionDate", "B-Currency", "I-Currency", "B-Total", "I-Total", "B-TotalTax", "I-TotalTax", "B-InvoiceNumber", "I-InvoiceNumber", "B-Subtotal", "I-Subtotal", "B-LineItems", "I-LineItems" ]
This repository contains:
microsoft/layoutlmv3-baseTo fine-tune this model on your custom data:
# On RunPod GPU pod or local machine with GPU
python main.py --mode train --push-to-hub --version v1.0
This will:
from transformers import LayoutLMv3ForTokenClassification, LayoutLMv3Processor
from PIL import Image
# Load model and processor
model = LayoutLMv3ForTokenClassification.from_pretrained("mkdigitalgmbh/runpo-LayoutLM3-Invoice-Receipt")
processor = LayoutLMv3Processor.from_pretrained("mkdigitalgmbh/runpo-LayoutLM3-Invoice-Receipt", apply_ocr=False)
# Prepare inputs (you need OCR results: words and bounding boxes)
image = Image.open("receipt.jpg").convert("RGB")
words = ["STORE", "NAME", "Total:", "$10.99"]
boxes = [[10, 10, 100, 30], [110, 10, 200, 30], [10, 50, 80, 70], [90, 50, 150, 70]]
# Normalize boxes to 0-1000 range
width, height = image.size
normalized_boxes = [[int(1000*x0/width), int(1000*y0/height),
int(1000*x1/width), int(1000*y1/height)] for x0,y0,x1,y1 in boxes]
encoding = processor(image, words, boxes=normalized_boxes, return_tensors="pt")
outputs = model(**encoding)
predictions = outputs.logits.argmax(-1)
This model is designed for deployment on RunPod Serverless:
Build and push Docker image:
cd deployment/runpod/LayoutLMv3
python deploy.py --action deploy
Create RunPod endpoint:
registry.hf.space/your-username/layoutlmv3-inference:latestHF_REPO_ID=mkdigitalgmbh/runpo-LayoutLM3-Invoice-ReceiptHF_TOKEN=<your-token>MODEL_VERSION=main (or specific version tag after training)The model uses IOB (Inside-Outside-Beginning) tagging:
Text: ["Total:", "$", "10", ".", "99"]
Labels: ["B-Total", "I-Total", "I-Total", "I-Total", "I-Total"]
Extracted: Total: "$ 10 . 99"
| Version | Date | Description | Status |
|---|---|---|---|
| main | 2025-11-13 | Initialized with base model + custom labels | Base (not trained) |
After training, versions will be tagged (v1.0, v1.1, etc.).
When training is performed, the following configuration will be used:
{
"model_name": "microsoft/layoutlmv3-base",
"learning_rate": 5e-05,
"batch_size": 4,
"num_epochs": 20,
"warmup_steps": 500,
"max_length": 512,
"validation_split": 0.2,
"random_seed": 42,
"gradient_accumulation_steps": 2,
"eval_steps": 100,
"save_steps": 500,
"logging_steps": 50
}
@misc{layoutlmv3-receipt-invoice,
author = {MK Digital GmbH},
title = {LayoutLMv3 Receipt/Invoice Field Extraction},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/mkdigitalgmbh/runpo-LayoutLM3-Invoice-Receipt}}
}
@article{huang2022layoutlmv3,
title={LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking},
author={Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu},
journal={arXiv preprint arXiv:2204.08387},
year={2022}
}
Apache 2.0
For questions or issues, please open an issue in the repository.