Downloads · 30 days
0
shubham-Bgs/Text-normalization-hindi
Text-normalization-hindi is a machine learning model from shubham-Bgs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is fine-tuned for text normalization in Hindi. It converts non-standard entities—such as dates, currencies, and scientific units—into their fully normalized forms.
Downloads · 30 days
0
Access
Public
Updated Feb 9, 2025
Repo size
242 MB
Likes
1
Public
Click a slice to open those files.
.pth242 MB · 100%
From the Hugging Face model README
This model is fine-tuned for text normalization in Hindi. It converts non-standard entities—such as dates, currencies, and scientific units—into their fully normalized forms.
T5-smallSPRINGLab/IndicVoices-R_Hindi, enriched with synthetic examples for dates, currencies, and units.2e-532 (with gradient accumulation)You can use the model with the transformers library:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("shubham-Bgs/Text-normalization-hindi")
model = AutoModelForSeq2SeqLM.from_pretrained("shubham-Bgs/Text-normalization-hindi")
# Example input
input_text = "15 / 03 / 1990 को, वैज्ञानिक ने $120 में 500 mg का नमूना खरीदा।"
inputs = tokenizer(input_text, return_tensors="pt", padding=True)
# Generate normalized text
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))