Downloads · 30 days
0
Mer1Alii/TR-ECommerce-CustomerSupport-Tokenizer
TR-ECommerce-CustomerSupport-Tokenizer is a machine learning model from Mer1Alii. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A custom-trained Byte-Pair Encoding (BPE) tokenizer optimized specifically for Turkish e-commerce customer support dialogues. Trained on the Mer1Alii/TR-ECommerce-CustomerSupport-Instructions corpus, this tokenizer dr…
Downloads · 30 days
0
Access
Public
Updated Jul 18, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json415 KB · 99%
From the Hugging Face model README
A custom-trained Byte-Pair Encoding (BPE) tokenizer optimized specifically for Turkish e-commerce customer support dialogues. Trained on the Mer1Alii/TR-ECommerce-CustomerSupport-Instructions corpus, this tokenizer drastically improves token efficiency and semantic comprehension for Turkish conversational AI.
Turkish is an agglutinative language with a rich morphological structure. Words are constructed by attaching multiple suffixes to a root (e.g., kar-go-lar-ı-mız-dan).
Standard English-centric tokenizers (like GPT-2 or LLaMA) do not have these Turkish roots/suffixes in their pre-trained vocabularies. As a result, they fragment basic Turkish words into tiny, meaningless character groups. This leads to:
This custom tokenizer solves these issues by learning a vocabulary derived directly from real Turkish customer service dialogues.
Here is a comparison of how different tokenizers split the sample Turkish e-commerce query:
"kargom teslim edilmedi iade istiyorum" (my package was not delivered, I want a return)
| Tokenizer | Tokenized Representation | Token Count | Efficiency Gain |
|---|---|---|---|
| GPT-2 (Standard) | ['k', 'arg', 'om', ' t', 'es', 'lim', ' ed', 'il', 'medi', ' i', 'ade', ' is', 't', 'iy', 'orum'] | 15 | Baseline |
| Our Custom Tokenizer | ['kargom', ' teslim', ' edil', 'medi', ' iade', ' istiyorum'] | 6 | 2.5x Fewer Tokens (60% Savings) |
ByteLevelBPETokenizer)Mer1Alii/TR-ECommerce-CustomerSupport-Instructions)<s>: Beginning of Sequence (BOS)<pad>: Padding (PAD)</s>: End of Sequence (EOS)<unk>: Unknown token (UNK)<mask>: Masking token (MASK)You can load and use this tokenizer directly in Python using the Hugging Face transformers library:
from transformers import AutoTokenizer
# Load custom tokenizer
tokenizer = AutoTokenizer.from_pretrained("Mer1Alii/TR-ECommerce-CustomerSupport-Tokenizer")
# Test Sentence
text = "kargom teslim edilmedi iade istiyorum"
tokens = tokenizer.encode(text)
print("Token IDs:", tokens)
print("Decoded Tokens:", tokenizer.convert_ids_to_tokens(tokens))