Downloads · 30 days
15
1% of all-time downloads
MoseliMotsoehli/zuBERTa
zuBERTa is a fill-mask model from MoseliMotsoehli. Use it when you need the model to fill a missing word. It is set up for transformers.
zuBERTa is a RoBERTa style transformer language model trained on zulu text.
Downloads · 30 days
15
1% of all-time downloads
All-time downloads
2K
Public
Repo size
1.2 GB
Likes
1
Public
Click a slice to open those files.
.h5375 MB · 41%
From the Hugging Face model README
zuBERTa is a RoBERTa style transformer language model trained on zulu text.
The model can be used for getting embeddings to use on a down-stream task such as question answering.
>>> from transformers import pipeline
>>> from transformers import AutoTokenizer, AutoModelWithLMHead
>>> tokenizer = AutoTokenizer.from_pretrained("MoseliMotsoehli/zuBERTa")
>>> model = AutoModelWithLMHead.from_pretrained("MoseliMotsoehli/zuBERTa")
>>> unmasker = pipeline('fill-mask', model=model, tokenizer=tokenizer)
>>> unmasker("Abafika eNkandla bafika sebeholwa <mask> uMpongo kaZingelwayo.")
[
{
"sequence": "<s>Abafika eNkandla bafika sebeholwa khona uMpongo kaZingelwayo.</s>",
"score": 0.050459690392017365,
"token": 555,
"token_str": "Ġkhona"
},
{
"sequence": "<s>Abafika eNkandla bafika sebeholwa inkosi uMpongo kaZingelwayo.</s>",
"score": 0.03668094798922539,
"token": 2321,
"token_str": "Ġinkosi"
},
{
"sequence": "<s>Abafika eNkandla bafika sebeholwa ubukhosi uMpongo kaZingelwayo.</s>",
"score": 0.028774697333574295,
"token": 5101,
"token_str": "Ġubukhosi"
}
]
@inproceedings{author = {Moseli Motsoehli},
title = {Towards transformation of Southern African language models through transformers.},
year={2020}
}