Downloads · 30 days
0
yasserius/bangla-tokenizer
bangla-tokenizer is a machine learning model from yasserius. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A specialized Bangla/Bengali tokenizer extracted from multilingual models and extended with missing characters.
Downloads · 30 days
0
Access
Public
Updated Aug 31, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.4 MB · 78%
From the Hugging Face model README
A specialized Bangla/Bengali tokenizer extracted from multilingual models and extended with missing characters.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yasserius/bangla-tokenizer")
# Tokenize Bangla text
text = "আমি বাংলায় কথা বলি"
tokens = tokenizer.tokenize(text)
print(tokens)