Downloads · 30 days
0
Johnnyman1100/EZ-Tokenizer_The_Tokenizer
EZ-Tokenizer_The_Tokenizer is a machine learning model from Johnnyman1100. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
"Go ahead, try to break it. I dare you." - A tokenizer so efficient, it feels like cheating.
Downloads · 30 days
0
Access
Public
Updated May 30, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json3.6 MB · 100%
From the Hugging Face model README
"Go ahead, try to break it. I dare you." - A tokenizer so efficient, it feels like cheating.
from tokenizers import Tokenizer
tokenizer = Tokenizer.from_pretrained("johnnyman1100/EZ-Tokenizer_The_Tokenizer")
# Test it yourself
text = "Your text here"
encoded = tokenizer.encode(text)
decoded = tokenizer.decode(encoded.ids)
assert text == decoded # Try to make this fail, I'll wait...
print(f"Compression: {len(text)/len(encoded.ids):.2f} chars/token")
Find any text where this tokenizer:
First to report a verified case gets a shoutout!
Because in a world of bloated models, efficiency still wins. This tokenizer proves you don't need 100K+ tokens to achieve perfect reconstruction and better compression.
MIT
"I didn't believe it either until I saw the benchmarks." - You, probably