Downloads · 30 days
0
gorkemergune/my-tokenizer
my-tokenizer is a text generation model from gorkemergune. Use it when you need the model to write or continue text. It is set up for tokenizers. The card lists the license as mit.
byte-level Byte-Pair Encoding (BPE) tokenizer trained from scratch with the tokenizers library. It was built as a learning exercise: text is scraped from a Wikipedia article, a BPE tokenizer is trained on it, and the…
Downloads · 30 days
0
Access
Public
Updated Jul 19, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json127 KB · 97%
From the Hugging Face model README
byte-level Byte-Pair Encoding (BPE) tokenizer trained from scratch with
the tokenizers library. It was built as a learning exercise: text is scraped
from a Wikipedia article, a BPE tokenizer is trained on it, and the result is
wrapped as a PreTrainedTokenizerFast so it loads through AutoTokenizer.
| Property | Value |
|---|---|
| Algorithm | Byte-level BPE |
| Vocabulary size | 2048 (2¹¹) |
| Special tokens | <unk>, <pad>, <bos>, <eos> |
| Pre-tokenizer | ByteLevel(add_prefix_space=False) |
| Decoder | ByteLevel |
| Training data | Plain text of the English Wikipedia article Large language model |
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("gorkemergune/my-tokenizer")
example = "Hello! This is a very small tokenizer example."
token_ids = tokenizer.encode(example)
print("Tokens:", tokenizer.convert_ids_to_tokens(token_ids))
print("Token IDs:", token_ids)
print("Decoded:", tokenizer.decode(token_ids))
The training pipeline is available on GitHub (gorkemergune/wiki2bpe). In short:
scraper.py downloads the Wikipedia article as plain text into text.txt.script.py trains the byte-level BPE tokenizer on text.txt and pushes it here.