Downloads · 30 days
14
48% of all-time downloads
Pradhap1125/t5-small-sentence-validator
t5-small-sentence-validator is a machine learning model from Pradhap1125. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is the model card of a 🤗 Transformers model fine-tuned for text normalization — restoring punctuation, capitalization, and proper spacing from raw transcripts or unformatted text.
Downloads · 30 days
14
48% of all-time downloads
All-time downloads
29
Public
Parameters
60.5M
969 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt484 MB · 66%
From the Hugging Face model README
This is the model card of a 🤗 Transformers model fine-tuned for text normalization — restoring punctuation, capitalization, and proper spacing from raw transcripts or unformatted text.
Pradhap Rajamani
Funded by : Independent academic project — Purdue University (Master’s in Computer Science)
Shared by : Pradhap Rajamani
t5-smallAutomatically adds missing punctuation, capitalization, and spacing to unformatted English text.
Example:
Input: normalize: helloeveryonewelcome to todayssession
Output: Hello everyone, welcome to today's session.
Text cleanup for ASR and transcripts
Preprocessing for NLP tasks (NER, summarization, etc.)
Chatbot or conversation logs normalization
🚫 Out-of-Scope Use
Rewriting semantics or paraphrasing
Limitation Description
| Limitation | Description |
|---|---|
| Ambiguous punctuation | May insert or omit commas/periods differently from human editors |
| Case sensitivity | Rarely fails to capitalize proper nouns |
| Domain limitation | Trained primarily on general English text (Wikipedia-style) |
Always perform a quick review before deploying results in production systems.
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_name = "pradhap1125/t5-small-sentence-validator"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
text = "helloeveryonewelcome to todayssession"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
A custom English dataset derived from Wikipedia sentences. Text was noisified by removing punctuation, capitalization, and spacing to simulate ASR output.
| Field | Description |
|---|---|
| Source | English Wikipedia corpus |
| Preprocessing | Removed punctuation, altered casing, and spacing |
| Training Samples | ~90,000 |
| Validation Samples | ~2,5000 |
| Parameter | Value |
|---|---|
| Model | T5-small |
| Epochs | 3 |
| Batch Size | 16 |
| Learning Rate | 3e-4 |
| Weight Decay | 0.01 |
| Optimizer | AdamW |
| Scheduler | Linear decay |
| Loss Function | CrossEntropyLoss |
| Evaluation Metric | Token-level accuracy, loss |
| Setting | Specification |
|---|---|
| Hardware | NVIDIA Tesla T4 GPU |
| Platform | Google Colab |
| Framework | PyTorch |
| Library | Hugging Face Transformers |
| Runtime | ~4 hours |
| Metric | Value |
|---|---|
| eval_loss | 0.0744 |
| Input | Output |
|---|---|
| helloeveryonewelcome to todayssession | Hello everyone, welcome to today's session. |
| this isatest ofthe normalizationmodel | This is a test of the normalization model. |
| itwasagoodday today | It was a good day today. |
Pradhap Rajamani
📧 https://www.linkedin.com/in/pradhap-rajamani/ 📧 https://github.com/pradhap1125