Downloads · 30 days
0
sidskarki/phi4-nepali-tokenizer
phi4-nepali-tokenizer is a machine learning model from sidskarki. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Extended tokenizer for Phi-4 with ~15K added high-value Nepali/Devanagari tokens.
Downloads · 30 days
0
Access
Public
Updated May 12, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json10.2 MB · 100%
From the Hugging Face model README
Extended tokenizer for Phi-4 with ~15K added high-value Nepali/Devanagari tokens.
| Nepali tok/word | |
|---|---|
| Original Phi-4 | 7.10 |
| Extended (this) | 3.41 |
| Reduction | 51.9% |
The extended tokenizer is a drop-in replacement for the original. To use the new tokens effectively, the model needs continued pretraining on Nepali text (see the Qwen3-4B Nepali model for a full CPT+SFT example).
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("sidskarki/phi4-nepali-tokenizer")
tokens = tokenizer.tokenize("नेपालको राजधानी काठमाडौं हो")
print(tokens, len(tokens))
Part of a 17-model Nepali tokenizer benchmark measuring the Nepali token tax across modern LLM tokenizers.