Downloads · 30 days
27
5% of all-time downloads
daviddrzik/SK_BPE_BLM
SK_BPE_BLM is a fill-mask model from daviddrzik. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
SKBPEBLM is a pretrained small language model for the Slovak language, based on the RoBERTa architecture. The model utilizes standard Byte-Pair Encoding (BPE) tokenization (pureBPE, more info here) and is case-insensi…
Downloads · 30 days
27
5% of all-time downloads
All-time downloads
573
Public
Parameters
58.7M
235 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors235 MB · 99%
From the Hugging Face model README
SK_BPE_BLM is a pretrained small language model for the Slovak language, based on the RoBERTa architecture. The model utilizes standard Byte-Pair Encoding (BPE) tokenization (pureBPE, more info here) and is case-insensitive, meaning it operates in lowercase. While the pretrained model can be used for masked language modeling, it is primarily intended for fine-tuning on downstream NLP tasks.
To use the SK_BPE_BLM model, follow these steps:
from transformers import pipeline, RobertaTokenizer, AutoModelForMaskedLM
# Load the custom tokenizer and model
tokenizer = RobertaTokenizer.from_pretrained("daviddrzik/SK_BPE_BLM")
model = AutoModelForMaskedLM.from_pretrained("daviddrzik/SK_BPE_BLM")
# Create a pipeline with the custom model and tokenizer
unmasker = pipeline('fill-mask', model=model, tokenizer=tokenizer)
# Use the pipeline
result = unmasker("včera večer sme <mask> nový film v kine, ktorý mal premiéru iba pred týždňom.")
print(result)
[{'score': 0.2665567100048065,
'token': 18599,
'token_str': ' pozreli',
'sequence': 'včera večer sme pozreli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
{'score': 0.23860174417495728,
'token': 1056,
'token_str': ' mali',
'sequence': 'včera večer sme mali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
{'score': 0.1962040513753891,
'token': 6915,
'token_str': ' videli',
'sequence': 'včera večer sme videli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
{'score': 0.03656836599111557,
'token': 26996,
'token_str': ' pozerali',
'sequence': 'včera večer sme pozerali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
{'score': 0.030735589563846588,
'token': 9058,
'token_str': ' objavili',
'sequence': 'včera večer sme objavili nový film v kine, ktorý mal premiéru iba pred týždňom.'}]
The SK_BPE_BLM model was pretrained using a subset of the OSCAR 2019 corpus, specifically focusing on the Slovak language. The corpus underwent comprehensive preprocessing to ensure the quality and relevance of the data:
Additionally, the preprocessing included further refinement steps to create the final dataset:
After preprocessing, the training corpus consisted of:
The SK_BPE_BLM model was trained with the following key parameters:
The model was trained using the Hugging Face library, but without using the Trainer class—native PyTorch was used instead.
Here are the fine-tuned versions of the SK_BPE_BLM model based on the folders provided:
SK_BPE_BLM-ner: Fine-tuned for Named Entity Recognition (NER) tasks.SK_BPE_BLM-pos: Fine-tuned for Part-of-Speech (POS) tagging.SK_BPE_BLM-qa: Fine-tuned for Question Answering tasks.SK_BPE_BLM-sentiment-csfd: Fine-tuned for sentiment analysis on the CSFD (movie review) dataset.SK_BPE_BLM-sentiment-multidomain: Fine-tuned for sentiment analysis across multiple domains.SK_BPE_BLM-sentiment-reviews: Fine-tuned for sentiment analysis on general review datasets.SK_BPE_BLM-topic-news: Fine-tuned for topic classification in news articles.If you find our model or paper useful, please consider citing our work:
Part of this research (specifically the models pre-trained for 10 epochs) has been published in:
Držík, D., & Forgac, F. (2024). Slovak morphological tokenizer using the Byte-Pair Encoding algorithm. PeerJ Computer Science, 10, e2465. https://doi.org/10.7717/peerj-cs.2465
An extended version of this work (including models pre-trained for 30 epochs and evaluated across multiple NLP downstream tasks) has been published in:
Držík, D., & Kapusta, J. (2026). The importance of morphology-aware subword tokenization for NLP tasks in Slovak language modeling. Expert Systems with Applications, 312, 131492. https://doi.org/10.1016/j.eswa.2026.131492
@article{drzik2024slovak,
title={Slovak morphological tokenizer using the Byte-Pair Encoding algorithm},
author={Držík, Dávid and Forgac, František},
journal={PeerJ Computer Science},
volume={10},
pages={e2465},
year={2024},
month={11},
issn={2376-5992},
doi={10.7717/peerj-cs.2465}
}
@article{Drzik2026,
author = {Dávid Držík and Jozef Kapusta},
title = {The importance of morphology-aware subword tokenization for NLP tasks in Slovak language modeling},
journal = {Expert Systems with Applications},
volume = {312},
pages = {131492},
year = {2026},
month = {5},
issn = {09574174},
doi = {10.1016/j.eswa.2026.131492}
}