Downloads · 30 days
1.3K
5% of all-time downloads
Eraly-ml/KazBERT
KazBERT is a fill-mask model from Eraly-ml. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
<details <summary<span style="color:4CAF50;"<strongLicense & Metadata</strong</span</summary
Downloads · 30 days
1.3K
5% of all-time downloads
All-time downloads
24.6K
Public
Parameters
111M
4.5 GB on disk
Likes
15
Public
Click a slice to open those files.
.json1.2 GB · 38%
From the Hugging Face model README
"KazBERT қазақ тілін [MASK] түсінеді."If you find KazBERT useful please press like button
KazBERT is a BERT-based model fine-tuned specifically for Kazakh using Masked Language Modeling (MLM). It is based on bert-base-uncased and uses a custom tokenizer trained on Kazakh text.
Erlanulu, Y. G. (2025). KazBERT: A Custom BERT Model for the Kazakh Language. Zenodo.
config.json – Model configmodel.safetensors – Model weightstokenizer.json – Tokenizer datatokenizer_config.json – Tokenizer configspecial_tokens_map.json – Special tokensvocab.txt – VocabularyInstall 🤗 Transformers and load the model:
from transformers import BertForMaskedLM, BertTokenizerFast
model_name = "Eraly-ml/KazBERT"
tokenizer = BertTokenizerFast.from_pretrained(model_name)
model = BertForMaskedLM.from_pretrained(model_name)
from transformers import pipeline
pipe = pipeline("fill-mask", model="Eraly-ml/KazBERT")
output = pipe('KazBERT қазақ тілін [MASK] түсінеді.')
Output:
[
{"score": 0.198, "token_str": "жетік", "sequence": "KazBERT қазақ тілін жетік түсінеді."},
{"score": 0.038, "token_str": "де", "sequence": "KazBERT қазақ тілін де түсінеді."},
{"score": 0.032, "token_str": "терең", "sequence": "KazBERT қазақ тілін терең түсінеді."},
{"score": 0.029, "token_str": "ерте", "sequence": "KazBERT қазақ тілін ерте түсінеді."},
{"score": 0.026, "token_str": "жете", "sequence": "KazBERT қазақ тілін жете түсінеді."}
]
- Trained only on public Kazakh Wikipedia & Common Crawl
- Might miss informal speech or dialects
- Could underperform on deep-context or rare words
- May reflect cultural or social biases in data
Apache 2.0 License
@misc{erlanulu_2025_15565394,
author = {Erlanulu, Yeraly Gainulla},
title = {KazBERT: A Custom BERT Model for the Kazakh
Language
},
month = may,
year = 2025,
publisher = {Zenodo},
doi = {10.5281/zenodo.15565394},
url = {https://doi.org/10.5281/zenodo.15565394},
}