Downloads · 30 days
94
55% of all-time downloads
AETHORIA-AI/TR-HASH-Tokenizer-32K
TR-HASH-Tokenizer-32K is a machine learning model from AETHORIA-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
The 32,000-token ByteLevel BPE tokenizer used by the TR-HASH language-model research line and the 200B-token pretraining mixture.
Downloads · 30 days
94
55% of all-time downloads
All-time downloads
170
Public
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.3 MB · 100%
From the Hugging Face model README
The 32,000-token ByteLevel BPE tokenizer used by the TR-HASH language-model research line and the 200B-token pretraining mixture.
</s> (ID 0)<pad> (ID 1)<s> (ID 2)<unk> (ID 3)from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
"AETHORIA-AI/TR-HASH-Tokenizer-32K"
)
encoded = tokenizer("Token routing starts with tokenization.")
decoded = tokenizer.decode(encoded["input_ids"])
This repository contains the tokenizer only. It does not define a chat template or a model architecture.
r50k_baser50k_base is used here as the approximately 50K-token GPT reference. The
comparison uses transformers without added special tokens for TR-HASH and
tiktoken for r50k_base.
| Encoding | Vocabulary | Tokens on fixed suite | Characters/token |
|---|---|---|---|
| TR-HASH Tokenizer 32K | 32,000 | 375 | 3.712 |
GPT r50k_base | 50,257 | 354 | 3.932 |
TR-HASH uses a 36.3% smaller vocabulary. On the fixed 1,392-character
suite included in benchmark_r50k.py, it produces 5.9%
more tokens than r50k_base. At hidden size 1,024, the smaller vocabulary
removes 18,695,168 embedding parameters, or about 37.4 MB in BF16/FP16, when
input and output embeddings are tied.
This small suite covers English prose, technical text, code/JSON, mathematics, and French. It is an illustrative, reproducible check rather than a claim about every corpus. Compression should be measured again on the intended training or deployment distribution.
pip install transformers tiktoken
python benchmark_r50k.py