Downloads · 30 days
0
Murjani/chessnano
chessnano is a text generation model from Murjani. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
A compact autoregressive transformer that reads chess moves in Standard Algebraic Notation and predicts the next one. No board representation and no evaluation function — the model sees notation and nothing else.
Downloads · 30 days
0
Access
Public
Updated Aug 19, 2026
Repo size
165 MB
Likes
1
Public
Click a slice to open those files.
.pt165 MB · 100%
From the Hugging Face model README
A compact autoregressive transformer that reads chess moves in Standard Algebraic Notation and predicts the next one. No board representation and no evaluation function — the model sees notation and nothing else.
Code: github.com/Kcbir/chessnano
| File | Size | Description |
|---|---|---|
chessnano_deployed.pt | 58.4 MB | Quantised artifact — INT8 weights, bit-packed sign bits, fp16 embedding |
chessnano_fp16.pt | 102 MB | fp16 checkpoint |
vocab.json | 76 KB | SAN token vocabulary |
| Layers | 8 |
| Model dim | 768 |
| Attention | grouped-query, 12 query heads / 4 KV heads, head dim 64 |
| Positions | RoPE |
| Feedforward | SwiGLU, 2048 hidden |
| Normalisation | RMSNorm, pre-norm, no bias |
| Vocabulary | 2048 SAN tokens |
| Context | 512 tokens |
| Parameters | 51.9M |
Weights are trained with quantization simulated in the forward pass. The TurboQuant path applies a Hadamard rotation, quantises each coordinate pair's angle to 128 levels, and stores a one-bit Johnson–Lindenstrauss residual sketch scaled by the least-squares coefficient. Training and export share one reconstruction function, so the simulated quantizer matches the exported one.
import chess
from inference.tot_inference import ChessNanoPlayer, TotConfig
player = ChessNanoPlayer("chessnano_deployed.pt", tot_cfg=TotConfig(depth=1))
board, history = chess.Board(), []
for san in ["e4", "e5", "Nf3"]:
board.push_san(san)
history.append(san)
print(player.predict_move(history, board))
Candidate moves are masked against python-chess's legal move list before
sampling, so the decoder cannot emit an illegal move.
MIT