Downloads · 30 days
1K
1% of all-time downloads
youscan/ukr-roberta-base
ukr-roberta-base is a fill-mask model from youscan. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
Below is the list of corpora used along with the output of wc command (counting lines, words and characters). These corpora were concatenated and tokenized with HuggingFace Roberta Tokenizer.
Downloads · 30 days
1K
1% of all-time downloads
All-time downloads
92.8K
Public
Repo size
2.5 GB
Likes
25
Public
Click a slice to open those files.
.bin507 MB · 50%
From the Hugging Face model README
Below is the list of corpora used along with the output of wc command (counting lines, words and characters). These corpora were concatenated and tokenized with HuggingFace Roberta Tokenizer.
| Tables | Lines | Words | Characters |
|---|---|---|---|
| Ukrainian Wikipedia - May 2020 | 18 001 466 | 201 207 739 | 2 647 891 947 |
| Ukrainian OSCAR deduplicated dataset | 56 560 011 | 2 250 210 650 | 29 705 050 592 |
| Sampled mentions from social networks | 11 245 710 | 128 461 796 | 1 632 567 763 |
| Total | 85 807 187 | 2 579 880 185 | 33 985 510 302 |
Vitalii Radchenko - contact me on Twitter @vitaliradchenko