Downloads · 30 days
14
17% of all-time downloads
olaverse/mist-encoder-base-ng
mist-encoder-base-ng is a fill-mask model from olaverse. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
14
17% of all-time downloads
All-time downloads
84
Public
Parameters
30.9M
124 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors124 MB · 97%
From the Hugging Face model README

A small (30.9M-parameter) modern encoder specialised for Nigerian languages —
Hausa (ha), Yoruba (yo), Igbo (ig), and Nigerian Pidgin (pcm) — pretrained from
scratch with a masked-language-modeling (MLM) objective using the unified
olaverse/otk-bpe-50k (Naija) tokenizer.
It is a deliberate specialist: a compact base you attach task heads to (classification, NER, language-ID, sentence embeddings). It is not intended to compete on raw task accuracy with larger multilingual or African-language encoders — its value is efficiency, a low-fertility Nigerian tokenizer, explicit Pidgin support, 0% UNK, and a clean Apache-2.0 release.
Load the encoder body and attach a head:
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("olaverse/mist-encoder-base-ng")
enc = AutoModel.from_pretrained("olaverse/mist-encoder-base-ng")
Good fits: topic/sentiment/language-ID classification, sentence embeddings (contrastive fine-tuning), and on-device / low-resource deployment where 28–30M params matters. NER is supported but weaker than larger models (see below).
All sources are commercial-friendly (attribution-only), consistent with the Apache-2.0 release:
| Source | License | Role |
|---|---|---|
| FineWeb-2 (ha/yo/ig/pcm) | ODC-By | Web text |
| castorini/wura (Nigerian subset) | Apache-2.0 | Audited mC4 + news |
| asr-nigerian-pidgin/nigerian-pidgin-1.0 | CC-BY-4.0 | Fresh Pidgin sentences |
FineWeb-2 and WURA both descend from Common Crawl / mC4, so documents were cross-deduped. The corpus was language-balanced (abundant Hausa capped; scarce Igbo/Pidgin taken in full, with the smallest language lightly upsampled) and chunked into 254-token windows so all text was used rather than truncating each document. Final training corpus: ~480k chunks.
olaverse/otk-bpe-50k unified Naija — byte-level BPE, ~50k vocab, 0% UNK,
NFC diacritic preservation, code-mixed English support.Three benchmarks, all four languages, compared against AfriBERTa (v2, 126M) and mBERT (178M). Numbers are honest and include where the model is weaker.
From the otk-bpe-50k unified-Naija benchmark (MasakhaNEWS):
| Tokenizer | Hausa | Yoruba | Igbo | Pidgin |
|---|---|---|---|---|
| otk-bpe-50k (ours) | 1.231 | 1.296 | 1.416 | 1.249 |
| GPT-4o (o200k) | 1.589 | 1.687 | 1.807 | 1.304 |
| AfroXLMR | 1.604 | 2.277 | 2.570 | 1.401 |
Lower fertility = more signal per token at a fixed sequence length. The tokenizer beats general multilingual tokenizers on all four languages.
| Model | Params | Hausa | Yoruba | Igbo | Pidgin |
|---|---|---|---|---|---|
| mist-encoder-base-ng | 30.9M | 0.878 | 0.859 | 0.803 | 0.898 |
| AfriBERTa | 126M | 0.924 | 0.921 | 0.914 | 0.991 |
| mBERT | 178M | 0.806 | 0.886 | 0.805 | 0.967 |
Competitive at a fraction of the size — beats mBERT on Hausa, ties on Igbo, trails AfriBERTa.
| Model | Params | Hausa | Yoruba | Igbo | Pidgin |
|---|---|---|---|---|---|
| mist-encoder-base-ng | 30.9M | 0.656 | 0.779 | 0.804 | 0.729 |
| AfriBERTa | 126M | 0.850 | 0.867 | 0.897 | 0.886 |
| mBERT | 178M | 0.810 | 0.837 | 0.855 | 0.881 |
The model trails both baselines on NER. This is the honest weak spot — see below.
Evidence-backed, in priority order: (1) more model capacity (~50–80M) — NER is where the parameter gap bit hardest; (2) a less-fragmenting tokenizer for token tasks (larger vocab or per-language merge budgets); (3) more pretraining data, especially Pidgin and Hausa. Longer pretraining is not a lever — eval loss already plateaued by ~epoch 11.
Apache-2.0. Training data is attribution-only (ODC-By / Apache-2.0 / CC-BY-4.0); please retain attribution to the upstream datasets.
Built with FineWeb-2, WURA (Oladipo et al., EMNLP 2023), and the Nigerian Pidgin ASR corpus. Evaluated on MasakhaNEWS and MasakhaNER 2.0 (Adelani et al., Masakhane). AfriBERTa (Ogueji et al., 2021) and mBERT (Devlin et al., 2019) used as comparison baselines.