Downloads · 30 days
0
Hailay/geez-en-shared-tokenizer
geez-en-shared-tokenizer is a machine learning model from Hailay. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as unknown.
A shared SentencePiece Unigram tokenizer (vocab size 120,000, bytefallback=true) covering Amharic, Tigrinya, Tigre, Classical Ge'ez, and English. Built as a continuation of "MoVoC: Morphology-Aware Subword Constructio…
Downloads · 30 days
0
Access
Public
Updated Jul 31, 2026
Repo size
2.8 MB
Likes
0
Public
Click a slice to open those files.
.model2.8 MB · 51%
From the Hugging Face model README
A shared SentencePiece Unigram tokenizer (vocab size 120,000,
byte_fallback=true) covering Amharic, Tigrinya, Tigre, Classical Ge'ez, and
English. Built as a continuation of
"MoVoC: Morphology-Aware Subword Construction for Ge'ez Script Languages"
(Teklehaymanot, Fazlija, Nejdl — L3S Research Center).
import sentencepiece as spm
sp = spm.SentencePieceProcessor(model_file="tokenizer.model")
print(sp.encode("አልሰበሩም", out_type=str))
Trained on a temperature-sampled (alpha=0.3) mix of cleaned text from these
sources. Licenses are mixed and one local source is unspecified — review
before treating this as uniformly open-licensed for downstream redistribution
of text generated by decoding, though the vocabulary/subword-piece list
itself does not reproduce the training text.
| Language | Primary sources | License(s) |
|---|---|---|
| Amharic | castorini/afriberta-corpus, michsethowusu/amharic-sentiments-corpus, yordanoswuletaw/amharic-pretraining-corpus, rasyosef/amharic-sentences-corpus, cis-lmu/Glot500, cis-lmu/GlotCC-V1, HuggingFaceFW/finetranslations | per HF dataset card / CC0-1.0 / ODC-By |
| Tigrinya | fgaim/GLOCR-Tigrinya, fgaim/tigrinya-squad, mewaeltsegay/TigrinyaLargeText, SIMBA9657/haddas-tigrinya-corpus, cis-lmu/Glot500/GlotCC-V1, HuggingFaceFW/finetranslations, local file ~/Ser/train.txt (1.98M lines) | per HF dataset card / CC0-1.0 / ODC-By / unspecified (local) |
| Tigre | BeitTigreAI/tigre-data-lexicon, BeitTigreAI/tigre-data-monolingual-text (primary, 490K rows) | CC-BY-SA-4.0 |
| Ge'ez | Bedru/Eng-Geez (2,107 pairs — the only real corpus found anywhere; see Limitations) | per HF dataset card |
| English | translated_text field from HuggingFaceFW/finetranslations | ODC-By |
Full per-source breakdown: see SOURCES.md in the
GeezTokenizer project this was built from.
byte_fallback=true lets this tokenizer
still encode Bilin text losslessly, but with no dedicated subword coverage.BeitTigreAI corpus —
automatic language-ID is known to confuse Tigre with the much more common
Tigrinya.0% UNK rate across all 5 languages; fertility 1.2–1.4 tokens/word. Full
report and methodology in the source project's 04_validation/report.md.
pip install tokenizers huggingface_hub
| Property | Value |
|---|---|
| Algorithm | SentencePiece Unigram |
| Vocabulary size | 120,000 |
| Byte fallback | enabled |
Character coverage, sampling temperature, normalisation rules, and the per-language corpus sizes used for temperature sampling are Not documented.
If you use this artifact, please cite:
@inproceedings{teklehaymanot2025movoc,
title = {MoVoC: Morphology-Aware Subword Construction for Ge'ez Script Languages},
author = {Teklehaymanot, Hailay Kidu and Fazlija, Dren and Nejdl, Wolfgang},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
year = {2025},
url = {https://arxiv.org/abs/2509.08812}
}