Downloads · 30 days
0
ansul90/hindi-bpe-tokenizer
hindi-bpe-tokenizer is a machine learning model from ansul90. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for custom. The card lists the license as mit.
A Byte Pair Encoding (BPE) tokenizer optimized for Hindi text using Devanagari script.
Downloads · 30 days
0
Access
Public
Updated Nov 7, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.txt1.5 MB · 98%
From the Hugging Face model README
A Byte Pair Encoding (BPE) tokenizer optimized for Hindi text using Devanagari script.
\u0900-\u097F)# Clone the repository
git clone https://huggingface.co/ansul90/hindi-bpe-tokenizer
cd hindi-bpe-tokenizer
# Install dependencies
pip install regex numpy
# Or with uv:
uv add regex numpy
⚠️ Note: This repository does not include the pre-trained model file (543MB). You need to train it once locally, which takes only ~30 seconds.
python train_bpe_simple.py
This will:
hindi_bpe_tokenizer.json (~543MB)from hindi_bpe_tokenizer import HindiBPETokenizer
# Load trained tokenizer
tokenizer = HindiBPETokenizer()
tokenizer.load('hindi_bpe_tokenizer.json')
# Encode Hindi text
text = "भारत एक महान देश है।"
tokens = tokenizer.encode(text)
print(f"Tokens: {tokens}")
# Decode back to text
decoded = tokenizer.decode(tokens)
print(f"Decoded: {decoded}")
# Get compression statistics
stats = tokenizer.get_compression_stats(text)
print(f"Compression ratio: {stats['compression_ratio']:.2f}X")
print(f"Original bytes: {stats['original_bytes']}")
print(f"Compressed tokens: {stats['compressed_tokens']}")
| Metric | Value |
|---|---|
| Vocabulary Size | 5,500 tokens |
| Compression Ratio | 6.52X (avg), 10.44X (best) |
| Decoding Accuracy | 100% |
| Training Corpus | 575K chars, 1.5MB |
| Training Time | ~30 seconds |
| Category | Original Bytes | Compressed Tokens | Compression Ratio |
|---|---|---|---|
| Space Mission | 204 | 31 | 6.58X |
| Cricket News | 146 | 27 | 5.41X |
| Science & Tech | 188 | 18 | 10.44X |
| Language | 123 | 18 | 6.83X |
| Education | 140 | 17 | 8.24X |
| Environment | 132 | 21 | 6.29X |
| Mixed Content | 125 | 34 | 3.68X |
| Long Sentence | 240 | 33 | 7.27X |
from hindi_bpe_tokenizer import HindiBPETokenizer
# Create tokenizer with custom vocabulary size
tokenizer = HindiBPETokenizer(vocab_size=8000)
# Load your custom Hindi corpus
with open('my_corpus.txt', 'r', encoding='utf-8') as f:
corpus = f.read()
# Train
tokenizer.train(corpus, verbose=True)
# Save
tokenizer.save('my_custom_tokenizer.json')
stats = tokenizer.get_compression_stats("हिंदी टेक्स्ट")
print(f"Original characters: {stats['original_chars']}")
print(f"Original bytes: {stats['original_bytes']}")
print(f"Compressed tokens: {stats['compressed_tokens']}")
print(f"Compression ratio: {stats['compression_ratio']:.2f}X")
print(f"Vocabulary size: {stats['vocab_size']:,}")
hindi-bpe-tokenizer/
├── hindi_bpe_tokenizer.py # Core implementation (8KB)
├── train_bpe_simple.py # Training script (5KB)
├── create_diverse_hindi_corpus.py # Corpus generator (17KB)
├── hindi_corpus.txt # Training data (1.5MB)
├── training_results.json # Performance metrics (2KB)
├── pyproject.toml # Dependencies
└── README.md # This file
Note: hindi_bpe_tokenizer.json (543MB) is generated when you run train_bpe_simple.py
The tokenizer was trained on diverse Hindi content including:
\u0900-\u097F, \u0980-\u09FFr""" ?[\u0900-\u097F]+| ?[\u0980-\u09FF]+| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+"""
regex library (for Unicode support)numpy (optional, for numerical operations)pip install regex
# Train the tokenizer first
python train_bpe_simple.py
# Ensure files are read with UTF-8 encoding
with open('file.txt', 'r', encoding='utf-8') as f:
content = f.read()
If you use this tokenizer in your research or project, please cite:
@misc{hindi_bpe_tokenizer_2025,
title={Hindi BPE Tokenizer: Byte Pair Encoding for Devanagari Script},
author={Your Name},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/ansul90/hindi-bpe-tokenizer}
}
MIT License - See LICENSE file for details
धन्यवाद (Thank you) for using Hindi BPE Tokenizer! 🙏