Downloads · 30 days
0
Talayhan/turkishwkpd-tokenizer
turkishwkpd-tokenizer is a machine learning model from Talayhan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This is a Byte-Pair Encoding (BPE) tokenizer trained from scratch on Turkish Wikipedia text using the 🤗 Hugging Face tokenizers library. It is intended for use as the tokenizer component of Turkish-language NLP models.
Downloads · 30 days
0
Access
Public
Updated Jul 18, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json1.2 MB · 100%
From the Hugging Face model README
This is a Byte-Pair Encoding (BPE) tokenizer trained from scratch on Turkish Wikipedia text using the 🤗 Hugging Face tokenizers library. It is intended for use as the tokenizer component of Turkish-language NLP models.
This is a Byte-Pair Encoding (BPE) tokenizer trained from scratch on Turkish Wikipedia text using the 🤗 Hugging Face tokenizers library. It is intended for use as the tokenizer component of Turkish-language NLP models.
This tokenizer can be used to convert Turkish text into subword tokens (BPE) for downstream use in training or fine-tuning Turkish NLP models (e.g., language models, classifiers).
Not recommended for languages other than Turkish, or for domains substantially different from Wikipedia text (e.g., code, social media slang, dialectal variants), where the vocabulary coverage may be poor.
As the tokenizer was trained solely on Turkish Wikipedia, it may underrepresent informal language, dialectal variation, code-switching, and domain-specific vocabulary (e.g., medical, legal, colloquial). A vocabulary size of 8,192 is relatively small, which may lead to more aggressive subword splitting for rare or out-of-domain words.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the tokenizer. Consider evaluating tokenization quality (fertility, coverage) on your target domain before use, especially if it differs from encyclopedic text.
Turkish Wikipedia dump, sourced from kaan39/turkish-wikipedia-dataset on the Hugging Face Hub. The dataset contains 26,871 rows with a total file size of 37.7 MB.
Only the messages (text content) field was extracted from the dataset and written to a plain .txt file; the tokenizer was trained on this extracted text.
BPE (Byte-Pair Encoding) subword tokenizer, vocabulary size 8,192, trained with the Hugging Face tokenizers library.