Downloads · 30 days
14
1% of all-time downloads
cfinley/punct_restore_fr
punct_restore_fr is a token classification model from cfinley. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
1% of all-time downloads
All-time downloads
971
Public
Repo size
2.6 GB
Likes
1
Public
Click a slice to open those files.
.bin440 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of camembert-base on a raw, French opensubtitles dataset. It achieves the following results on the evaluation set:
Classifies tokens based on beginning of French sentences (B-SENT) and everything else (O).
This model aims to help punctuation restoration on French YouTube auto-generated subtitles. In doing so, one can measure more in a corpus such as words per sentence, grammar structures per sentence, etc.
1 million Open Subtitles (French) sentences. 80%/10%/10% training/validation/test split.
The sentences:
Token/tag pairs batched together in groups of 64. This helps show variety of positions for B-SENT and O tags. This also keeps training examples from just being one sentence. Otherwise, this leads to having the first word and only the first word in a sequence being labeled B-SENT.
The following hyperparameters were used during training: