Downloads · 30 days
4.1K
2% of all-time downloads
EMBEDDIA/sloberta
sloberta is a fill-mask model from EMBEDDIA. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
Downloads · 30 days
4.1K
2% of all-time downloads
All-time downloads
230K
Public
Repo size
886 MB
Likes
5
Public
Click a slice to open those files.
.bin443 MB · 99%
From the Hugging Face model README
Load in transformers library with:
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/sloberta")
model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/sloberta")
SloBERTa model is a monolingual Slovene BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The corpora used for training the model have 3.47 billion tokens in total. The subword vocabulary contains 32,000 tokens. The scripts and programs used for data preparation and training the model are available on https://github.com/clarinsi/Slovene-BERT-Tool
SloBERTa was trained for 200,000 iterations or about 98 epochs.
The following corpora were used for training the model: