Downloads · 30 days
0
LeoSavi/Chess-God-Transformer
Chess-God-Transformer is a machine learning model from LeoSavi. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model is a custom-built Encoder-Decoder Transformer designed to predict the optimal next move in a chess game given a Board State in FEN (Forsyth-Edwards Notation).
Downloads · 30 days
0
Access
Public
Updated Mar 8, 2026
Repo size
267 MB
Likes
0
Public
Click a slice to open those files.
.pth44.5 MB · 99%
From the Hugging Face model README
This model is a custom-built Encoder-Decoder Transformer designed to predict the optimal next move in a chess game given a Board State in FEN (Forsyth-Edwards Notation).
This project was developed for my transformer class. I decided to challenge myself, rather than training an existing model I built it from scratch. terrible decision...but here we are.
I merged data from multiple sources, while paing attention to the limitation of my hardware and the needs of my AI.
bonna46/Chess-FEN-and-NL-Format. Used In both Fine tuning and Training.ssingh22/chess-evaluations (tactics subset). The data were split in two parts:
lichess/chess-puzzles. The challenge with these data was to unpack the moves, once done it the whole dataset size increased exponentially. Given the size of the dataset, in lychess_puzzles.py, I filter for themes I thought my AI lacked, moreover I took only the most highest rated and played moves. Ultimately I sampled for ~400000+ and got ~860801 data points.The core architecture is based on the original "Attention is All You Need" paper, specifically following the implementation guide from DataCamp's Transformer Tutorial.
ATTEMPT: Hyperparameters Initially I used Optuna for fine tuning the hyperparameters. I run 70 trials with 15 epochs each. We sampled 10% the training data, and used 80% for training and 20% for validation. The search algorithm, it is in similar fashion another project I did, and it focused on minimizing two factors:
NOTE: due to vram limitation of my GPU, I manually set hyperparameters and made ad-hoc changes to the architecture.
d_model: 256
num_heads: 8
num_layers: 6
d_ff: 1024
dropout: 0.1
lr: 0.0003
batch_size: 64
I chose these parameters because they represent the best balance between model capacity and the VRAM constraints of an RTX 4060 laptop (8GB).
The total parameter count is 11,086,884.
Given VRAM issues I tweaked the training and the architecture of the model as follow:

ssingh22/chess-evaluations tactics dataset and bonna46/Chess-FEN-and-NL-Format-30K-DatasetThe default Temperature after running 100 matches vs every stockfish-model (weak,mid,strong,GM). I calculated win-rate as 1pt. for Win, 0.5 draw and -1 for a loss. Codes are in tester.py. I decided that 0.65 is the optimal default temperature for winrate consistency across different runs and opponents. Below the last run I tried.

The training of the base model took about 9h30min, i.e. 20min per epoch * 28 epochs (because of early stopping). While finetuning about 30 min. the model TransformerGodPlayer.pth is saved in model folder and uploaded in HuggingFace along with the hyperparameters opt-configs.yml.
Libraries used are described in requirements.txt. If you want to install them in bulk you can run the following command once cd into the directory:
pip install -r requirements.txt
This model is ment to be used with the TransformerPlayer class in player.py after cloning the original repo.
git clone https://github.com/LeonSavi/chess_exam
For the purpose of the class tournment, the model automatically imports a ad-hoc ChessTokenizer and the Transformer class to load the .pth weights, first attemps to load those locally (since the GitHub repo is going to be cloned), otherwise import from HuggingFace.
from player import TransformerPlayer
model = TransformerPlayer() #everything is already initialized
fen = "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1"
move = model.get_move(fen)
print(f"God-Transformer predicts: {move}")
Fallback As a probabilistic model, the Transformer occasionally predicts illegal moves. particularly in unusual positions that differ from the training distribution. A python-chess validation layer is applied inside get_move() to catch these cases before they reach the game engine. The fallback strategy works in three stages: