Downloads · 30 days
9
2% of all-time downloads
nbroad/fix_punct_uncased_t5_small
fix_punct_uncased_t5_small is a machine learning model from nbroad. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of google/t5-v11-small on the NPR utterances dataset.
Downloads · 30 days
9
2% of all-time downloads
All-time downloads
399
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.bin308 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of google/t5-v1_1-small on the NPR utterances dataset.
The model was trained on 80k rows from the above dataset consisting of NPR radio transcripts. Commans, periods, and semicolons were removed from the text and then random commas, periods, and semicolons were added. The model was trained to place those three punctuation marks in the correct location. All texts were lowercase during training.
It achieves the following results on the evaluation set:
The purpose of this model is to correct the punctuation in a sentence. For example, the phrase "this is, a sentence. with odd punctuation to show off what, the model. can do" gets changed to "this is a sentence with odd punctuation to show off what the model can do."
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rouge1 | Rouge2 | Rougel | Rougelsum | Gen Len |
|---|---|---|---|---|---|---|---|---|
| 1.3066 | 1.0 | 600 | 0.4347 | 59.0002 | 54.7692 | 58.7112 | 58.7856 | 16.3808 |
| 0.8192 | 2.0 | 1200 | 0.3154 | 62.4672 | 59.0199 | 62.4096 | 62.3667 | 16.5158 |
| 0.7208 | 3.0 | 1800 | 0.3050 | 62.701 | 59.3201 | 62.6739 | 62.6165 | 16.5471 |