Downloads · 30 days
244
2% of all-time downloads
keshan/SinhalaBERTo
SinhalaBERTo is a fill-mask model from keshan. Use it when you need the model to fill a missing word. It is set up for transformers.
This is a slightly smaller model trained on OSCAR Sinhala dedup dataset. As Sinhala is one of those low resource languages, there are only a handful of models been trained. So, this would be a great place to start tra…
Downloads · 30 days
244
2% of all-time downloads
All-time downloads
10.7K
Public
Parameters
83.5M
2.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.pt668 MB · 31%
How the weights are stored.
F3283.5M · 100%
From the Hugging Face model README
This is a slightly smaller model trained on OSCAR Sinhala dedup dataset. As Sinhala is one of those low resource languages, there are only a handful of models been trained. So, this would be a great place to start training for more downstream tasks.
The model chosen for training is Roberta with the following specifications:
You can use this model directly with a pipeline for masked language modeling:
from transformers import AutoTokenizer, AutoModelWithLMHead, pipeline
model = AutoModelWithLMHead.from_pretrained("keshan/SinhalaBERTo")
tokenizer = AutoTokenizer.from_pretrained("keshan/SinhalaBERTo")
fill_mask = pipeline('fill-mask', model=model, tokenizer=tokenizer)
fill_mask("මම ගෙදර <mask>.")