Downloads · 30 days
0
Ankit91/hindi-bpe-tokenizer
hindi-bpe-tokenizer is a machine learning model from Ankit91. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This tokenizer was trained on a mixed Hindi dataset using Byte-Pair Encoding (BPE). - Vocabulary size: 6000 - Base tokens: UTF-8 bytes (256) - Trained in Python using a custom implementation - Original length: 7053637…
Downloads · 30 days
0
Access
Public
Updated Nov 7, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json693 KB · 89%
From the Hugging Face model README
This tokenizer was trained on a mixed Hindi dataset using Byte-Pair Encoding (BPE).
vocab.json: Token IDsmerges.json: Merge rulesmetadata.json: Tokenizer configuration