Downloads · 30 days
41
36% of all-time downloads
manaspros/indian-receipt-parser-v2
indian-receipt-parser-v2 is a text generation model from manaspros. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.1.
v2.0 is a comprehensive LoRA-finetuned version of Llama 3.1 8B specialized in parsing Indian financial receipts and invoices. This version represents a 5x improvement over v1.0 with massive expansion in vendor coverag…
Downloads · 30 days
41
36% of all-time downloads
All-time downloads
113
Public
Repo size
185 MB
Likes
2
Public
Click a slice to open those files.
.safetensors168 MB · 91%
From the Hugging Face model README
v2.0 is a comprehensive LoRA-finetuned version of Llama 3.1 8B specialized in parsing Indian financial receipts and invoices. This version represents a 5x improvement over v1.0 with massive expansion in vendor coverage and regional utilities.
Electricity Boards: BESCOM (Bengaluru), MSEDCL (Maharashtra), TANGEDCO (Tamil Nadu), Adani Electricity (Mumbai), Tata Power (Delhi/Mumbai), BSES Rajdhani/Yamuna (Delhi), UPPCL (UP), TSSPDCL (Telangana), CESC (Kolkata), BEST (Mumbai), KSEBL (Kerala), PSPCL (Punjab)
Gas Companies: IGL (Delhi NCR), MGL (Mumbai), Gujarat Gas, Adani Total Gas
Water Authorities: DJB (Delhi), BWSSB (Bengaluru), HMWSSB (Hyderabad), KWA (Kerala), CMWSSB (Chennai)
Regional Transport: APSRTC, TGSRTC, GSRTC, RSRTC, UPSRTC, KSRTC (Kerala & Karnataka), HRTC, WBSTC, MSRTC
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
# Load model
base_model = "meta-llama/Meta-Llama-3.1-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(base_model)
model = AutoModelForCausalLM.from_pretrained(
base_model,
device_map="auto",
torch_dtype=torch.bfloat16
)
model = PeftModel.from_pretrained(model, "manaspros/indian-receipt-parser-v2")
# Prepare receipt
receipt_text = """
Thank you for your purchase from Zomato
Date: November 15, 2025
Subtotal: ₹450
Tax (18%): ₹81
Total: ₹531
"""
messages = [
{
"role": "system",
"content": "You are a financial receipt parser specialized in Indian vendors. Extract vendor name, amount, date, and category from receipts. Return ONLY valid JSON."
},
{
"role": "user",
"content": receipt_text
}
]
# Generate
input_text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.1)
response = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(response)
# Output: {"vendor": "Zomato", "amount": 531.0, "date": "2025-11-15", "category": "dining", ...}
{
"vendor": "VendorName",
"amount": 0.0,
"date": "YYYY-MM-DD",
"category": "category",
"tax": 0.0,
"currency": "INR",
"confidence": 0.95
}
| Metric | v1.0 | v2.0 | Improvement |
|---|---|---|---|
| Vendors | 51 | 227 | +345% |
| Examples | 2,000 | 10,400 | +420% |
| Categories | 8 | 19 | +138% |
| Regional Coverage | None | 46 vendors | New |
| Final Loss | 0.11 | ~0.06 | 45% better |
| Training Time | 25 min | 4-5 hours | - |
Author: Manas Choudhary Repository: GitHub License: Llama 3.1 License Version: 2.0 Release Date: November 2025
@misc{indian-receipt-parser-v2,
author = {Manas Choudhary},
title = {Indian Receipt Parser v2.0: Comprehensive Receipt Parsing for Indian Vendors},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/manaspros/indian-receipt-parser-v2}}
}
🎉 Ready for production use! This model provides the most comprehensive Indian receipt parsing available.