Downloads · 30 days
0
tachiwin/tokenizer_64k
tokenizer_64k is a machine learning model from tachiwin. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tokenizers.
A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json12.3 MB · 99%
From the Hugging Face model README
A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.
| Component | Target |
|---|---|
| Modern + old exotic-language data | 70% |
| English | 10% |
| Spanish | 10% |
| Code | 10% |
The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.
Tokenizer.train_from_iterator over a streaming
line generator (constant memory regardless of corpus size)Total materialized corpus: 457,300,912 bytes (0.426 GiB)
The tokenizer was evaluated against the Tachiwin language catalogue. Languages without text samples are skipped. For each language, all available samples are concatenated ONLY within that language for aggregate fertility statistics. Metrics: characters/token, tokens/character, UTF-8 bytes/token, tokens/UTF-8 byte, exact round-trip preservation.
The Hugging Face BPE trainer does not expose an internal resumable
merge-state checkpoint. The recipe therefore treats the completed
tokenizer.json as the training checkpoint:
recipe/;tokenizer.json already exists, subsequent runs skip BPE training;