Downloads · 30 days
22
0% of all-time downloads
jrahn/ROOK-CLF-9m
ROOK-CLF-9m is a text classification model from jrahn. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
A 9M parameter chess move prediction model using a classification approach, reproducing Google DeepMind's "Grandmaster-Level Chess Without Search".
Downloads · 30 days
22
0% of all-time downloads
All-time downloads
7.1K
Public
Parameters
8.9M
53.6 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.9 MB · 99%
From the Hugging Face model README
A 9M parameter chess move prediction model using a classification approach, reproducing Google DeepMind's "Grandmaster-Level Chess Without Search".
ROOK-CLF-9M reproduces one specific ablation from the appendix of Ruoss et al. 2024 "Grandmaster-Level Chess Without Search": the 9M parameter model configuration trained on behavior cloning (action prediction only).
What is Reproduced:
What is Different:
Overview:
The model can be used for:
The model is not suitable for:
Reported in the LAION research note:
The model uses a custom tokenization scheme critical for proper inference:
Step 1: FEN Processing (77 characters fixed)
# Original FEN
fen = "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1"
# Process FEN to fixed 77-character format:
# 1. Expand numbers to dots (e.g., "8" → "........")
# 2. Remove slashes
# 3. Pad castling to 4 chars, en passant to 2 chars, halfmove to 3 chars, fullmove to 3 chars
def process_fen(fen):
position, turn, castling, en_passant, halfmove, fullmove = fen.split(" ")
# Expand empty squares: "8" → "........"
position = re.sub(r'\d+', lambda m: "." * int(m.group()), position)
position = position.replace("/", "") # Remove row separators
castling = castling.ljust(4, ".") # Pad to 4 chars
en_passant = en_passant.ljust(2, ".") # Pad to 2 chars
halfmove = halfmove.ljust(2, ".") + "." # Pad to 3 chars total
fullmove = fullmove.ljust(3, ".") # Pad to 3 chars
return "".join([position, turn, castling, en_passant, halfmove, fullmove])
# Result: exactly 77 characters
processed = process_fen(fen)
# "rnbqkbnrpppppppp................................PPPPPPPPRNBQKBNRwKQkq-...0..1.."
Step 2: Add [CLS] token and convert to token IDs
# Add classification token
final_input = processed + "[CLS]" # 78 characters total
# Convert to token IDs (character-level tokenization)
tokens = [char_to_id[c] for c in final_input] # 78 tokens
Complete example:
Input FEN: "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1"
Processed: "rnbqkbnrpppppppp................................PPPPPPPPRNBQKBNRwKQkq-...0..1.."
With [CLS]: "rnbqkbnrpppppppp................................PPPPPPPPRNBQKBNRwKQkq-...0..1..[CLS]"
Token IDs: [13, 11, 3, 12, 10, 3, 11, 13, 15, 15, 15, 15, 15, 15, 15, 15, ...] # 78 tokens
For in-browser inference, the model is exported to ONNX format:
# ONNX export for web deployment
import torch
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained("jrahn/ROOK-CLF-9m")
dummy_input = torch.randint(0, 88, (1, 78))
torch.onnx.export(
model,
dummy_input,
"rook_clf_9m.onnx",
input_names=['input_ids'],
output_names=['logits'],
dynamic_axes={'input_ids': {0: 'batch_size'}}
)
If you use this model, please cite both our work and the original paper:
@article{rook2024,
title={ROOK: Strategic Reasoning in Chess Without Search},
author={Rahn, Jonathan and Jitsev, Jenia and Sun, Qi},
journal={LAION Research Notes},
year={2024},
url={https://laion.ai/notes/rook/}
}
@article{ruoss2024grandmaster,
title={Grandmaster-level chess without search},
author={Ruoss, Anian and Delétang, Grégoire and McAleese, Nell and Genewein, Tim and Weidinger, Laura and Cai, Matteo and Weber, Théophane and Hutter, Marcus and Legg, Shane},
journal={arXiv preprint arXiv:2402.04494},
year={2024}
}
Jonathan Rahn - GitHub | Research Page
LAION research note: https://laion.ai/notes/rook/