Downloads · 30 days
27
1% of all-time downloads
pandafm/donut-es
donut-es is a image-text-to-text model from pandafm. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This is a fine-tuned version of the Donut architecture, specifically tailored for parsing retail receipts. Donut is a transformer-based model designed for document understanding, and it performs OCR-free parsing by di…
Downloads · 30 days
27
1% of all-time downloads
All-time downloads
2.3K
Public
Parameters
201M
20.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors806 MB · 99%
How the weights are stored.
F32201M · 100%
From the Hugging Face model README
This is a fine-tuned version of the Donut architecture, specifically tailored for parsing retail receipts. Donut is a transformer-based model designed for document understanding, and it performs OCR-free parsing by directly processing images into structured JSON outputs. This implementation was fine-tuned using a custom dataset of artificial and real receipts.
This model is intended to be used for parsing receipts into structured data, extracting information such as item names, quantities, prices, taxes, and total amounts directly from image inputs.
The model was trained on a mixture of synthetic and real-world receipts:
The artificial receipts were generated using a combination of background images, fonts, and custom templates to mimic real-world conditions, ensuring the model can handle various types of distortions such as noise, wrinkles, and lighting changes. The real receipts were annotated manually using a custom tool based on the Marimo app, which allowed for structured annotation of receipt elements.
The model was tested on both validation and test datasets, achieving the following results:
This model is available on Hugging Face and can be used as follows:
from transformers import DonutProcessor, VisionEncoderDecoderModel
from PIL import Image
import json
import torch
import re
# Load model and processor
print("Loading Donut model...")
processor = DonutProcessor.from_pretrained("pandafm/donut-es")
model = VisionEncoderDecoderModel.from_pretrained("pandafm/donut-es")
if torch.cuda.is_available():
device = torch.device("cuda")
model.to(device)
else:
model.encoder.to(torch.bfloat16)
print("Donut model loaded.")
# Open image of a receipt
image = Image.open("path_to_receipt_image.jpg")
# Process image and generate JSON output
pixel_values = processor(image, return_tensors="pt").pixel_values
if torch.cuda.is_available():
pixel_values = pixel_values.to(device)
else:
pixel_values = pixel_values.to(torch.bfloat16)
# Convert output to JSON
task_prompt = "<s_cord-v2>"
decoder_input_ids = processor.tokenizer(task_prompt, add_special_tokens=False, return_tensors="pt").input_ids
decoder_input_ids = decoder_input_ids.to(device)
# autoregressively generate sequence
result = model.generate(
pixel_values,
decoder_input_ids=decoder_input_ids,
max_length=model.decoder.config.max_position_embeddings,
pad_token_id=processor.tokenizer.pad_token_id,
eos_token_id=processor.tokenizer.eos_token_id,
bad_words_ids=[[processor.tokenizer.unk_token_id]],
return_dict_in_generate=True,
)
seq = processor.batch_decode(result.sequences)[0]
seq = seq.replace(processor.tokenizer.eos_token, "").replace(processor.tokenizer.pad_token, "")
seq = re.sub(r"<.*?>", "", seq, count=1).strip() # remove first task start token
seq = processor.token2json(seq)
This model was fine-tuned as part of a research project for a Bachelor's Degree, leveraging the Donut architecture and integrating tools like OpenCV for data generation. The final dataset included both synthetic and real-world receipts to improve robustness in parsing.
@thesis{pandafm2024DonutES, author = {David Florez Mazuera}, title = {Ticket Parser}, school = {Universidad de Murcia}, year = {2024}, address = {Murcia, España}, month = {June}, type = {Bachelor's thesis}, note = {Gines García Mateos}, url = {}, keywords = {donut, transformers, fine-tune}, }