Downloads · 30 days
0
smcproject/malayalam-bpe-tokenizer
malayalam-bpe-tokenizer is a machine learning model from smcproject. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tokenizers. The card lists the license as mit.
A Byte Pair Encoding (BPE) tokenizer trained on Malayalam text corpus. Trained using the HuggingFace tokenizers library with Metaspace pre-tokenization and NFC normalization for correct handling of Malayalam Unicode c…
Downloads · 30 days
0
Access
Public
Updated Feb 21, 2026
Repo size
—
Likes
3
Public
Click a slice to open those files.
.json1.1 MB · 100%
From the Hugging Face model README
A Byte Pair Encoding (BPE) tokenizer trained on Malayalam text corpus. Trained using the HuggingFace tokenizers library with Metaspace pre-tokenization and NFC normalization for correct handling of Malayalam Unicode conjuncts.
| Property | Value |
|---|---|
| Algorithm | BPE (Byte Pair Encoding) |
| Vocabulary size | 16,000 |
| Pre-tokenizer | Metaspace (▁) |
| Normalizer | NFC + Strip |
| Special tokens | <s>, </s>, <unk>, <pad>, <mask> |
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("smc/malayalam-bpe-tokenizer")
text = "മലയാളം ഒരു ദ്രാവിഡ ഭാഷയാണ്"
tokens = tokenizer.tokenize(text)
print(tokens)
encoded = tokenizer(text, return_tensors="pt")
print(encoded)