Downloads · 30 days
0
aakashMeghwar01/SindhiLM-Tokenizer-v3
SindhiLM-Tokenizer-v3 is a machine learning model from aakashMeghwar01. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This tokenizer was specifically engineered and empirically tested for the Sindhi language, transitioning from a Byte-Level BPE model to a Unigram algorithm with Whitespace pre-tokenization to respect Sindhi morphologi…
Downloads · 30 days
0
Access
Public
Updated Jun 19, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.2 MB · 100%
From the Hugging Face model README
This tokenizer was specifically engineered and empirically tested for the Sindhi language, transitioning from a Byte-Level BPE model to a Unigram algorithm with Whitespace pre-tokenization to respect Sindhi morphological boundaries and prevent the shattering of Perso-Arabic text.
Trained on 200000 documents from a newly filtered version of the Sindhi corpus. Filtering was upgraded from a heuristic script-ratio check to a FastText LangID (lid.176) pass to aggressively remove Farsi, Urdu, and Pashto contamination.
1.619 tokens per word (Tested on fresh, held-out documents).1/9 exact matches against native-speaker evaluations.گھ (gh) aspirated digraph is SOMETIMES split by this tokenizer.The following hand-corrected morpheme boundary splits failed during testing. This implies ambiguity requiring POS context unavailable to a pure subword tokenizer:
['گھر', 'ن'], but the tokenizer split into ['گ', 'ھ', 'ر', 'ن'].['ڇوڪر', 'ي'], but the tokenizer split into ['ڇوڪري'].['استاد', 'ان'], but the tokenizer split into ['استادن'].['لکھ', 'ندو'], but the tokenizer split into ['ل', 'ک', 'ھندو'].['پڙھ', 'يل'], but the tokenizer split into ['پڙهيل'].['وڃ', 'ي'], but the tokenizer split into ['وڃي'].['ننڍ', 'ڙو'], but the tokenizer split into ['ننڍڙو'].['لکھ', 'ند', 'ڙ'], but the tokenizer split into ['لکندڙ'].