Downloads · 30 days
0
fortik11/Tokenizer-Ky-Ru
Tokenizer-Ky-Ru is a text generation model from fortik11. Use it when you need the model to write or continue text. The card lists the license as cc-by-4.0.
BPEF - is a family of ultra-optimized, custom Byte-Level BPE tokenizers available in two configurations: a massive flagship 120,000 token vocabulary and a lean, compute-efficient 65,000 token vocabulary.
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
13.3 MB
Likes
0
Public
Click a slice to open those files.
.json18.2 MB · 100%
From the Hugging Face model README
BPEF - is a family of ultra-optimized, custom Byte-Level BPE tokenizers available in two configurations: a massive flagship 120,000 token vocabulary and a lean, compute-efficient 65,000 token vocabulary.
Engineered from scratch by fortik11, this suite is specifically designed to eliminate the heavy "tokenization tax" on Cyrillic and highly agglutinative Central Asian languages, redefining compression standards for Russian and Kyrgyz NLP.
While global mainstream tokenizers (OpenAI GPT-4, Google Gemma, Meta Llama) yield a poor density of only 2.2 — 3.5 bytes/token on regional languages, the Fortaki suite achieves up to 3x higher efficiency, packing complex structures into record-low token footprints.
| Language & Text Style | 🌟 120K Flagship (Bytes/Token) | 🔥 65K Junior (Bytes/Token) |
|---|---|---|
| Kyrgyz (Modern / Official & Tech) | ~10.50 | 9.93 |
| Kyrgyz (Complex Literary / Philosophical) | 8.30 | 7.33 |
| Russian (Complex Literary / Archaic) | 8.53 | 7.61 |
| Russian (Modern / Official & Tech) | ~8.80 | 7.72 |
Note: Even when halving the vocabulary from 120K to 65K, the Junior model retains over 88% of the flagship's compression efficiency due to a highly optimized BPE-merge heuristic strategy.
'пространство', 'человеческого'), your inference latency drops dramatically.\xd3) almost entirely. The vocabulary beautifully captures the inflected system of Russian and the deep agglutinative suffix chains of Kyrgyz ('сотвор' + 'ённое').You can call either version of the Fortaki tokenizer directly inside any modern pipeline:
from transformers import PreTrainedTokenizerFast
# Load your preferred version (120K or 65K)
tokenizer = PreTrainedTokenizerFast.from_pretrained("fortik11/<your-repository-name>")
text = "2026-жылы мамлекеттик кызматтарды санариптештирүү..."
encoded = tokenizer.encode(text)
print(f"Total tokens produced: {len(encoded)}")
This suite is released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. You are completely free to share, adapt, and use these tokenizers commercially, provided that appropriate credit is given to the original author: fortik11 (Fortaki).