Downloads · 30 days
15
3% of all-time downloads
arnolfokam/roberta-base-kin
roberta-base-kin is a token classification model from arnolfokam. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
roberta-base-kin is a model based on the fine-tuned RoBERTa base model. It has been trained to recognize four types of entities:
Downloads · 30 days
15
3% of all-time downloads
All-time downloads
599
Public
Repo size
1.4 GB
Likes
0
Public
Click a slice to open those files.
.bin496 MB · 99%
From the Hugging Face model README
roberta-base-kin is a model based on the fine-tuned RoBERTa base model. It has been trained to recognize four types of entities:
This model was fine-tuned on the Kinyarwanda corpus (kin) of the MasakhaNER dataset. However, we thresholded the number of entity groups per sentence in this dataset to 10 entity groups.
This model was trained on a single NVIDIA P5000 from Paperspace
We evaluated this model on the test split of the Kinyarwandan corpus (kin) present in the MasakhaNER with no thresholding.
| Model Name | Precision | Recall | F1-score |
|---|---|---|---|
| roberta-base-kin | 76.26 | 80.58 | 78.36 |
from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline
tokenizer = AutoTokenizer.from_pretrained("arnolfokam/roberta-base-kin")
model = AutoModelForTokenClassification.from_pretrained("arnolfokam/roberta-base-kin")
nlp = pipeline("ner", model=model, tokenizer=tokenizer)
example = "Rayon Sports yasinyishije rutahizamu w’Umurundi"
ner_results = nlp(example)
print(ner_results)