Downloads · 30 days
20
25% of all-time downloads
abnuel/yoruba_task1_punctuation_model
yoruba_task1_punctuation_model is a token classification model from abnuel. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
A BERT-based token classification model for punctuation restoration in Yoruba text. Given raw unpunctuated Yoruba text, this model predicts punctuation marks at the token level — a key preprocessing step for downstrea…
Downloads · 30 days
20
25% of all-time downloads
All-time downloads
80
Public
Parameters
177M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors709 MB · 99%
From the Hugging Face model README
A BERT-based token classification model for punctuation restoration in Yoruba text. Given raw unpunctuated Yoruba text, this model predicts punctuation marks at the token level — a key preprocessing step for downstream NLP tasks on this low-resource West African language.
📄 Dataset: abnuel/yor_punctuation (1M–10M tokens)
Yoruba is a tonal language spoken by approximately 40–50 million people, primarily in southwestern Nigeria. Despite its large speaker base, it remains significantly underrepresented in NLP tooling. Unpunctuated Yoruba text is common in real-world sources (social media, transcribed speech, scanned documents), creating a major barrier for parsing, translation, and other NLP tasks.
This model addresses that gap by restoring punctuation as a sequence labeling task, fine-tuned on a large Yoruba text corpus.
yo)The model predicts one of the following token-level labels:
| Label | Description |
|---|---|
O | No punctuation |
COMMA | Comma (,) |
PERIOD | Full stop (.) |
QUESTION | Question mark (?) |
EXCLAMATION | Exclamation mark (!) |
(Exact label set may vary — check config.json for the full id2label mapping.)
from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline
model_id = "abnuel/yoruba_task1_punctuation_model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
nlp = pipeline("token-classification", model=model, tokenizer=tokenizer)
text = "Mo lọ sí ọjà lánàánì mo ra ẹran àti ẹyin"
result = nlp(text)
print(result)
@misc{adegunlehin2025yoruba-punct,
author = {Abayomi Adegunlehin},
title = {Yoruba Punctuation Restoration Model},
year = {2025},
url = {https://huggingface.co/abnuel/yoruba_task1_punctuation_model}
}