Downloads · 30 days
67
26% of all-time downloads
TobiasLogic/ChessAggro
ChessAggro is a text generation model from TobiasLogic. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A small GPT-2 style language model built to play in the mlabonne Chess LLM Arena (https://huggingface.co/spaces/mlabonne/chessllm) and to hold its own against the strongest entry there, FlameF0X/ChessSLM.
Downloads · 30 days
67
26% of all-time downloads
All-time downloads
258
Public
Parameters
49.3M
197 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors197 MB · 98%
From the Hugging Face model README
A small GPT-2 style language model built to play in the mlabonne Chess LLM Arena (https://huggingface.co/spaces/mlabonne/chessllm) and to hold its own against the strongest entry there, FlameF0X/ChessSLM.
Every move, the arena hands the model the constant prompt "1." and nothing else. It never sees the board or the move history. It then constrains generation to the current legal moves in SAN (with check and mate markers stripped) and samples one of them at temperature 1.0. The model plays blind, so its whole policy is a fixed distribution over move shapes, renormalised to whatever is legal in the position.
Because the model cannot calculate, the only thing that matters is that fixed distribution. A passive, positionally correct distribution loses in this setting. A distribution that reliably reaches decisive, tactical positions and converts them into checkmates does much better.
Training used the mlabonne/chessllm dataset, which is the same source ChessSLM learned from. Those are roughly 1500 to 3000 Elo games that are full of quick tactical checkmates and sharp attacking play, which is exactly the decisiveness that wins blind games.
The recipe:
Measured with the exact arena mechanism (constant "1." prompt, token by token constrained sampling over legal moves at temperature 1.0). The benchmark script and the game record are included in this repository.
Reference point for how hard the opponent is: under the same mechanism ChessSLM beats a random mover about 66 percent, and it beat every strong chess trained model we tried in the 33 to 40 percent range.
Against FlameF0X/ChessSLM:
300 games: 39 wins, 44 losses, 217 draws, score 49.2 percent, 95 percent CI 43.5 to 54.8 percent. A separate 120 game run scored 50.8 percent (14 wins, 12 losses, 94 draws). Taken together the model sits right at parity with ChessSLM, a genuine coin flip against the strongest entry in the arena. The proof.pgn file in this repository is the full game record from the 300 game run.
The decisive game balance is roughly even, which is the real change. Earlier attempts that trained on strong 2400 plus games lost the decisive games badly, around three to one, because sound but passive play gets mated in the blind setting. Reaching parity with the top arena entry is the result here rather than a significant win, and closing the last gap would take either a sharper edge or a much larger number of games to resolve.
pip install torch transformers python-chess
python benchmark.py TobiasLogic/ChessAggro FlameF0X/ChessSLM 200 out.pgn
The script is the arena mechanism itself, so the numbers are auditable.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("TobiasLogic/ChessAggro")
model = AutoModelForCausalLM.from_pretrained("TobiasLogic/ChessAggro")
Standard GPT-2 tokenizer and architecture, so it loads as a drop in causal LM.