Downloads · 30 days
52
1% of all-time downloads
jean-paul/KinyaBERT-large
KinyaBERT-large is a fill-mask model from jean-paul. Use it when you need the model to fill a missing word. It is set up for transformers.
A Pretrained model on the Kinyarwanda language dataset using a masked language modeling (MLM) objective. The BERT model was first introduced in this paper. This KinyaBERT model was pretrained with uncased tokens which…
Downloads · 30 days
52
1% of all-time downloads
All-time downloads
5K
Public
Parameters
109M
873 MB on disk
Likes
1
Public
Click a slice to open those files.
.bin437 MB · 50%
How the weights are stored.
F32109M · 100%
From the Hugging Face model README
A Pretrained model on the Kinyarwanda language dataset using a masked language modeling (MLM) objective. The BERT model was first introduced in this paper. This KinyaBERT model was pretrained with uncased tokens which means that no difference between for example ikinyarwanda and Ikinyarwanda.
The data set used has both sources from the new articles in Rwanda extracted from different new web pages, dumped Wikipedia files, and the books in Kinyarwanda. The sizes of the sources of data are 72 thousand new articles, three thousand dumped Wikipedia articles, and six books with more than a thousand pages.
The model was trained with the default configuration of BERT and Trainer from the Huggingface. However, due to some resource computation issues, we kept the number of transformer layers to 12.
from transformers import pipeline
the_mask_pipe = pipeline(
"fill-mask",
model='jean-paul/KinyaBERT-large',
tokenizer='jean-paul/KinyaBERT-large',
)
the_mask_pipe("Ejo ndikwiga nagize [MASK] baje kunsura.")
[{'sequence': 'ejo ndikwiga nagize amahirwe baje kunsura.', 'score': 0.3704017996788025, 'token': 1501, 'token_str': 'amahirwe'},
{'sequence': 'ejo ndikwiga nagize ngo baje kunsura.', 'score': 0.30745452642440796, 'token': 196, 'token_str': 'ngo'},
{'sequence': 'ejo ndikwiga nagize agahinda baje kunsura.', 'score': 0.0638100653886795, 'token': 3917, 'token_str': 'agahinda'},
{'sequence': 'ejo ndikwiga nagize ubwoba baje kunsura.', 'score': 0.04934622719883919, 'token': 2387, 'token_str': 'ubwoba'},
{'sequence': 'ejo ndikwiga nagizengo baje kunsura.', 'score': 0.02243402972817421, 'token': 455, 'token_str': '##ngo'}]
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("jean-paul/KinyaBERT-large")
model = AutoModelForMaskedLM.from_pretrained("jean-paul/KinyaBERT-large")
input_text = "Ejo ndikwiga nagize abashyitsi baje kunsura."
encoded_input = tokenizer(input_text, return_tensors='pt')
output = model(**encoded_input)
Note: We used the huggingface implementations for pretraining BERT from scratch, both the BERT model and the classes needed to do it.