Downloads · 30 days
0
PersianML/persian-bpe-tokenizer
persian-bpe-tokenizer is a text generation model from PersianML. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
The PersianBPETokenizer is a custom tokenizer specifically designed for the Persian (Farsi) language. It leverages the Byte-Pair Encoding (BPE) algorithm to create a robust vocabulary that can effectively handle the u…
Downloads · 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.3 MB · 99%
From the Hugging Face model README
The PersianBPETokenizer is a custom tokenizer specifically designed for the Persian (Farsi) language. It leverages the Byte-Pair Encoding (BPE) algorithm to create a robust vocabulary that can effectively handle the unique characteristics of Persian text. This tokenizer is optimized for use with advanced language models like BERT and RoBERTa, making it a valuable tool for various Persian NLP tasks.

If you use this tokenizer in your research, please cite it as:
Mohammad Shojaei. (2024). PersianBPETokenizer [Software]. Available at https://huggingface.co/mshojaei77/PersianBPETokenizer.
mshojaei77/PersianTelegramChannels[UNK]).datasets, tokenizers, transformers).mshojaei77/PersianTelegramChannels dataset and created a batch iterator for efficient training.PreTrainedTokenizerFast object for compatibility with Hugging Face Transformers.[UNK], [CLS], [SEP], [PAD], [MASK]To use the PersianBPETokenizer, first install the required libraries:
pip install -q --upgrade datasets tokenizers transformers
You can load the tokenizer using the Hugging Face Transformers library:
from transformers import AutoTokenizer
persian_tokenizer = AutoTokenizer.from_pretrained("mshojaei77/PersianBPETokenizer")
test_sentence = "سلام، چطور هستید؟ امیدوارم روز خوبی داشته باشید"
tokens = persian_tokenizer.tokenize(test_sentence)
print("Tokens:", tokens)
encoded = persian_tokenizer(test_sentence)
print("Input IDs:", encoded["input_ids"])
print("Decoded:", persian_tokenizer.decode(encoded["input_ids"]))
mshojaei77/PersianTelegramChannelsdatasets, tokenizers, and transformers