Downloads · 30 days
12
2% of all-time downloads
cjvt/sloberta-sleng
sloberta-sleng is a fill-mask model from cjvt. Use it when you need the model to fill a missing word. It is set up for transformers.
SloBERTa-SlEng is a masked language model, based on the SloBERTa Slovene model.
Downloads · 30 days
12
2% of all-time downloads
All-time downloads
710
Public
Repo size
467 MB
Likes
0
Public
Click a slice to open those files.
.bin234 MB · 99%
From the Hugging Face model README
SloBERTa-SlEng is a masked language model, based on the SloBERTa Slovene model.
SloBERTa-SlEng replaces the tokenizer, vocabulary and the embeddings layer of the SloBERTa model. The tokenizer and vocabulary used are bilingual, Slovene-English, based on conversational, non-standard, and slang language the model was trained on. They are the same as in the SlEng-bert model. The new embedding weights were initialized from the SloBERTa embeddings.
The new SloBERTa-SlEng model is SloBERTa model, which was further pre-trained for two epochs on the conversational English and Slovene corpora, the same as the SlEng-bert model.
The model was trained on English and Slovene tweets, Slovene corpora MaCoCu and Frenk, and a small subset of English Oscar corpus. We tried to keep the sizes of English and Slovene corpora as equal as possible. Training corpora had in total about 2.7 billion words.