Downloads · 30 days
557
8% of all-time downloads
EMBEDDIA/est-roberta
est-roberta is a fill-mask model from EMBEDDIA. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
Downloads · 30 days
557
8% of all-time downloads
All-time downloads
7K
Public
Repo size
936 MB
Likes
4
Public
Click a slice to open those files.
.bin467 MB · 99%
From the Hugging Face model README
Load in transformers library with:
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/est-roberta")
model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/est-roberta")
Est-RoBERTa model is a monolingual Estonian BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The Estonian corpora used for training the model have 2.51 billion tokens in total. The subword vocabulary contains 40,000 tokens.
Est-RoBERTa was trained for 40 epochs.