Downloads · 30 days
0
shwethd/kannada-tokenizer-50k
kannada-tokenizer-50k is a machine learning model from shwethd. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tokenizers. The card lists the license as mit.
A production-ready Byte Pair Encoding (BPE) tokenizer for Kannada language with 50,000 tokens.
Downloads · 30 days
0
Access
Public
Updated Nov 13, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json4.5 MB · 100%
From the Hugging Face model README
A production-ready Byte Pair Encoding (BPE) tokenizer for Kannada language with 50,000 tokens.
This tokenizer is specifically trained for the Kannada language using Wikipedia data. It achieves excellent compression ratios and handles Kannada morphology effectively through pure statistical learning.
pip install tokenizers
from tokenizers import Tokenizer
# Load the tokenizer
tokenizer = Tokenizer.from_file("tokenizer.json")
# Tokenize Kannada text
text = "ಕನ್ನಡ ಭಾಷೆಯು ಸುಂದರವಾಗಿದೆ"
encoding = tokenizer.encode(text)
print(f"Text: {text}")
print(f"Tokens: {encoding.tokens}")
print(f"IDs: {encoding.ids}")
# Decode back
decoded = tokenizer.decode(encoding.ids)
print(f"Decoded: {decoded}")
texts = [
"ಕನ್ನಡ ಭಾಷೆ",
"ಬೆಂಗಳೂರು ನಗರ",
"ಕರ್ನಾಟಕ ರಾಜ್ಯ"
]
encodings = tokenizer.encode_batch(texts)
for text, encoding in zip(texts, encodings):
print(f"{text} → {encoding.tokens}")
Systematic scaling study was conducted with vocabularies of 8K, 16K, 32K, 50K, 64K, and 100K. 50K was identified as optimal through:
| Vocabulary | Compression | Generalization Gap | Efficiency |
|---|---|---|---|
| 8,000 | 3.51 | 6.5% | baseline |
| 16,000 | 3.73 | - | 100% |
| 32,000 | 4.21 | 6.5% | 110% |
| 50,000 | 4.48 | 1.9% ⭐ | 55% |
| 64,000 | 4.62 | 7.4% | 35% |
| 100,000 | 4.81 | 13.1% | 24% |
50K achieves the best generalization with excellent compression!
Comprehensive evaluation on 9 different tests:
Overall Quality Score: 67% raw / 92% weighted (Production-Ready!)
| Tokenizer | Vocabulary | Type | Our Status |
|---|---|---|---|
| charanhu/kannada-tokenizer | 32,000 | Kannada-only | 1.56x larger |
| ruthuvikas1998/kannada-tokenizer | ~32-50K | Kannada-only | Comparable/larger |
| GPT-4 (multilingual) | ~100K total | Multilingual | Better for Kannada (specialized) |
This tokenizer is suitable for:
"ಕನ್ನಡ ಭಾಷೆ" → ['ಕನ್ನಡ', 'ಭಾಷೆ'] (2 tokens)
"ಬೆಂಗಳೂರು ನಗರ" → ['ಬೆಂಗಳೂರು', 'ನಗರ'] (2 tokens)
"ಮಗುವನ್ನು" → ['ಮಗುವನ್ನು'] (1 token)
"ಚಳಿಗಾಲ" → ['ಚಳಿಗಾಲ'] (1 token)
"ಮನೆಗೆ" → ['ಮನೆಗೆ'] (to house)
"ಮನೆಯಿಂದ" → ['ಮನೆಯಿಂದ'] (from house)
"ಮನೆಯಲ್ಲಿ" → ['ಮನೆಯಲ್ಲಿ'] (in house)
"ಕನ್ನಡ ದಕ್ಷಿಣ ಭಾರತದ ಕರ್ನಾಟಕ ರಾಜ್ಯದ ಅಧಿಕೃತ ಭಾಷೆಯಾಗಿದೆ"
→ 8 tokens, 4.6 chars/token compression
[PAD] (ID: 0) - Padding token[UNK] (ID: 1) - Unknown token[CLS] (ID: 2) - Classification token[SEP] (ID: 3) - Separator token[MASK] (ID: 4) - Mask token (for MLM tasks)Why Whitespace Pre-tokenizer?
Why 50K Vocabulary?
Why NFC Normalization?
MIT License - Free for commercial and academic use
If you use this tokenizer in your research, please cite:
@misc{kannada-bpe-tokenizer-2025,
title={Kannada BPE Tokenizer: Optimal Vocabulary Size Analysis},
author={shwethd},
year={2025},
note={50K-token BPE tokenizer trained on Kannada Wikipedia with systematic scaling analysis},
url={https://huggingface.co/shwethd/kannada-tokenizer}
}
Built with ❤️ for Kannada NLP