Downloads · 30 days
0
bench-labs/GCTokenizer-v1
GCTokenizer-v1 is a machine learning model from bench-labs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
GCT v1 is a corpus-free multilingual tokenizer built from vocabulary consensus across six open frontier tokenizers: DeepSeek, Kimi, Qwen, GLM, Gemma, and Mistral.
Downloads · 30 days
0
Access
Public
Updated Aug 11, 2026
Repo size
205 MB
Likes
13
Public
Click a slice to open those files.
.json238 MB · 100%
From the Hugging Face model README
GCT v1 is a corpus-free multilingual tokenizer built from vocabulary consensus across six open frontier tokenizers: DeepSeek, Kimi, Qwen, GLM, Gemma, and Mistral.
It provides four deterministic variants:
GCT prioritizes tokens independently selected by multiple existing tokenizers. This produces a strong shared Latin/technical vocabulary and universal UTF-8 byte fallback, while sacrificing some compression efficiency and multilingual balance compared with individually corpus-trained tokenizers.
The tokenizer uses no corpus-trained merges and performs exact lossless round-tripping in its reference implementation.
GCT is designed as a general-purpose tokenizer built from vocabulary consensus across multiple large language-model tokenizers.
To characterize its multilingual behavior, GCT was evaluated on 20 languages using the same measurement procedure for each language. The metric reported is tokens per byte; lower values indicate more compact tokenization.
Measurements below use GCT-L on Wikipedia plain-text extracts (~30 KB UTF-8 per language), encoded with greedy longest-match segmentation and byte fallback.
| Language | Tokens / Byte |
|---|---|
| English | 0.429 |
| Spanish | 0.476 |
| French | 0.472 |
| German | 0.438 |
| Italian | 0.455 |
| Portuguese | 0.487 |
| Romanian | 0.505 |
| Dutch | 0.454 |
| Polish | 0.515 |
| Czech | 0.561 |
| Russian | 0.973 |
| Ukrainian | 0.992 |
| Greek | 0.995 |
| Arabic | 0.995 |
| Hebrew | 0.994 |
| Hindi | 0.997 |
| Bengali | 0.996 |
| Chinese | 0.980 |
| Japanese | 0.992 |
| Korean | 0.985 |
These measurements are provided as a characterization of GCT's multilingual tokenization behavior, rather than as a comparison or leaderboard against other tokenizers.
Latin-script languages compress well under the shared consensus vocabulary. Non-Latin scripts fall back more often to UTF-8 byte tokens at the S/M/L tiers; GCT-XL recovers substantially more multi-byte pieces for several scripts (see multilingual_compression.json for full S/M/L/XL tables).
GCT includes deterministic byte fallback, allowing arbitrary byte sequences to be represented even when no corresponding vocabulary token exists.
The tokenizer has been tested against random byte sequences, arbitrary UTF-8, invalid UTF-8, control bytes, combining marks, emoji, repeated bytes, and all 256 possible byte values.
Across 40,000 randomized round-trip tests (10,000 per vocabulary tier), every input was reconstructed exactly:
decode(encode(x)) == x
40,000 / 40,000 — 100% exact recovery.
| Tier | Random cases | Failures | Exact recovery |
|---|---|---|---|
| S | 10,000 | 0 | 100% |
| M | 10,000 | 0 | 100% |
| L | 10,000 | 0 | 100% |
| XL | 10,000 | 0 | 100% |