Downloads · 30 days
0
Eraly-ml/KazBERT-benchmark
KazBERT-benchmark is a fill-mask model from Eraly-ml. Use it when you need the model to fill a missing word. The card lists the license as apache-2.0.
🤖 Fully AI-generated. Design, code (benchmark.py, benchmarkunt.py), plots and this card were produced end-to-end by an AI agent (Claude, Hermes ML research loop). Compute: a single Kaggle T4, inference-only, no train…
Downloads · 30 days
0
Access
Public
Updated Jul 5, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.png393 KB · 92%
From the Hugging Face model README
🤖 Fully AI-generated. Design, code (
benchmark.py,benchmark_unt.py), plots and this card were produced end-to-end by an AI agent (Claude, Hermes ML research loop). Compute: a single Kaggle T4, inference-only, no training. No credentials used or embedded — all models/datasets are public and the code is included for transparency.
ModernBERT-style comparison of KazBERT vs four Kazakh-capable encoders, zero-shot, across tokenizer efficiency, exam-QA accuracy and embedding quality.
dastur-mc reading MC, (2) the full Kazakh ЕНТ / UNT — 14,850 real exam questions across 7 subjects.
| model | vocab | fertility ↓ | MC-PLL ↑ | MC-embed ↑ | sec |
|---|---|---|---|---|---|
| KazBERT | 32k | 1.636 | 42.4% | 36.8% | 80.5 |
| mBERT | 119k | 2.759 | 27.5% | 45.8% | 271.6 |
| XLM-R base | 250k | 2.150 | 34.8% | 33.8% | 262.8 |
| kaz-roberta | 52k | 1.598 | 44.5% | 38.8% | 60.4 |
| KazakhBERTmulti | 100k | 1.457 | 36.6% | 40.5% | 95.6 |
(bold = best in column; random MC = 25%)
![]() | ![]() |
![]() | ![]() |
Астана — Қазақстанның [MASK] қаласы.
| model | top-3 |
|---|---|
| KazBERT | астана, ірі, алматы |
| mBERT | Астана, бар, 1 |
| XLM-R base | бас, 1, 19 |
Мен қазақ [MASK] сөйлеймін.
| model | top-3 |
|---|---|
| KazBERT | тілінде, тіліне, тілі |
| mBERT | ##қа, ##стан, ##та |
| XLM-R base | тілінде, ша, тілін |
Абай Құнанбаев — ұлы қазақ [MASK].
| model | top-3 |
|---|---|
| KazBERT | ақыны, энциклопедиясы, сср |
| mBERT | [UNK], ##ты, ##тар |
| XLM-R base | ақын, жазушы, ғалым |
KazBERT and the Kazakh-trained models return fluent Kazakh; mBERT often falls back to sub-word fragments or [UNK].
Zero-shot on every question of kazakh-unified-national-testing-mc (4–8 options each). Random baseline ≈ 17.9%.
| model | ЕНТ acc — PLL ↑ | ЕНТ acc — embed ↑ | sec |
|---|---|---|---|
| KazBERT | 22.3% | 21.3% | 801.8 |
| mBERT | 22.1% | 23.3% | 2631.3 |
| XLM-R base | 21.9% | 20.9% | 2867.7 |
| kaz-roberta | 22.7% | 24.2% | 680.2 |
| KazakhBERTmulti | 20.7% | 22.5% | 1173.4 |


Runs on a free Kaggle T4 with pip install transformers datasets — public models/datasets only:
python benchmark.py # Part 1: dastur-mc + tokenizer
python benchmark_unt.py # Part 2: full ЕНТ
kazakh-unified-national-testing-mc, kazakh-dastur-mc, kazakh_wiki_articlesml-research-loop