Downloads · 30 days
0
ankanmbz/chess-tok
chess-tok is a machine learning model from ankanmbz. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This chess tokenizer uses a large vocabulary (~844 tokens) with semantically meaningful units like 'w.', 'b.', piece+square combinations ('♙e4', '♞f6'), and complete suffixes ('..', '.x.', '.+').
Downloads · 30 days
0
Access
Public
Updated Jan 22, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json13.4 KB · 52%
From the Hugging Face model README
This chess tokenizer uses a large vocabulary (~844 tokens) with semantically meaningful units like 'w.', 'b.', piece+square combinations ('♙e4', '♞f6'), and complete suffixes ('..', '.x.', '.+').
This design reduces sequence length by ~60% compared to character-level tokenization, enabling faster training and better gradient flow in recurrent models.
from transformers import AutoTokenizer
# Load tokenizer directly from HuggingFace
tokenizer = AutoTokenizer.from_pretrained("ankanmbz/chess-tok", trust_remote_code=True)
# Tokenize chess moves
text = "w.♙e2♙e4.."
encoded = tokenizer(text, return_tensors="pt")
print(encoded)
# Decode
decoded = tokenizer.decode(encoded['input_ids'][0])
print(decoded)
# Batch processing
moves = ["w.♙e2♙e4..", "b.♟c7♟c5..", "w.♘g1♘f3.."]
batch = tokenizer(moves, padding=True, return_tensors="pt")
print(batch)
<pad>, <sos>, <eos>, <unk>844 tokens
w., b.♙e2, ♞f6, etc.a1, e4, h8, etc..., .x., .+, .+#, etc.Input: "w.♙e2♙e4.."
Tokens: ['w.', '♙e2', '♙e4', '..']
Token count: 4 (vs ~10 with character-level)
Apache 2.0