Downloads · 30 days
0
shatayumk/chesslm
chesslm is a text generation model from shatayumk. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
A from-scratch, Qwen3-style decoder-only transformer trained on Lichess games, built as a reimplementation of the 50M configuration from Understanding Reasoning from Pretraining to Post-Training (Table 2).
Downloads · 30 days
0
Access
Public
Updated Sep 14, 2026
Repo size
756 MB
Likes
0
Public
Click a slice to open those files.
.safetensors756 MB · 100%
From the Hugging Face model README
A from-scratch, Qwen3-style decoder-only transformer trained on Lichess games, built as a reimplementation of the 50M configuration from Understanding Reasoning from Pretraining to Post-Training (Table 2).
This repository holds four pretrained checkpoints. Three of them form a token-scaling sweep at fixed model size — same architecture, same data pipeline, only the token budget changes — so they can be compared directly against each other.
| Folder | Tokens seen | Opt. steps | Val loss | Embeddings |
|---|---|---|---|---|
50m_0.69b | 686,817,280 | 1,310 | 0.5319 | untied |
50m_1.1b | 1,149,763,584 | 2,193 | 0.4841 | untied |
50m_2.3b | 2,289,565,696 | 4,367 | 0.4537 | untied |
50m | 2,000,158,720 | 3,815 | not recorded | tied |
Validation loss is token-level cross-entropy on a held-out 20,000-game split (5,338 sequences of length 1024). These figures were reproduced from the released weights on a 128-sequence subsample of that split (0.5203 / 0.4844 / 0.4533), matching the recorded values to within sampling noise.
Note on the
50mfolder. This was the earlier run and it ties the input and output embeddings (tie_word_embeddings: true, 47,243,776 params, nolm_head), whereas the three sweep runs untie them (47,285,760 params). It also has no recorded validation loss. Treat it as a separate model, not as a fourth point on the scaling curve.
For the scaling sweep, loss falls monotonically as the token budget grows — 0.5319 → 0.4841 → 0.4537 for roughly 0.69B → 1.15B → 2.29B tokens.
| Parameters | 47.3M |
| Layers | 12 |
| Hidden size | 512 |
| Intermediate size | 1536 |
| Attention heads | 8 query / 4 key-value (GQA) |
| Head dim | 128 |
| Context length | 1024 |
| Vocab size | 82 |
| Position encoding | RoPE, θ = 1,000,000 |
| Normalization | RMSNorm (ε = 1e-6), incl. per-head Q/K norm |
| FFN | SwiGLU |
| Precision | fp32 |
Chess-specific, not a text tokenizer. Every move is exactly 4 tokens:
⟨piece⟩ ⟨from-square⟩ ⟨to-square⟩ ⟨flag⟩
The 82-token vocabulary is 6 piece symbols (P N B R Q K) + 64 square names (a1…h8) +
11 flags (- quiet, x capture, + check, # mate, =Q/=R/=B/=N promotions, O-O /
O-O-O castling, e.p.) + <EOS>.
modeling_chesslm.py and chess_tokenizer.py are included in this repo.
import json, torch, dataclasses
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
repo = snapshot_download("shatayumk/chesslm")
import sys; sys.path.insert(0, repo)
from modeling_chesslm import ChessLM, ModelConfig
from chess_tokenizer import ChessTokenizer
run = "50m_2.3b" # best of the sweep
cfg = json.load(open(f"{repo}/{run}/config.json"))
fields = {f.name for f in dataclasses.fields(ModelConfig)}
model = ChessLM(ModelConfig(**{k: v for k, v in cfg.items() if k in fields}))
model.load_state_dict(load_file(f"{repo}/{run}/model.safetensors"))
model.eval()
# Score a game
tok = ChessTokenizer()
ids = tok.encode_game(["e2e4", "e7e5", "g1f3"])
logits = model(torch.tensor(ids).unsqueeze(0))
print(logits.shape) # (1, 12, 82)
config.json carries the architecture fields plus a training block
(opt_step, tokens_seen, val_loss). The sweep folders also include the original
recipe.json. Note that ModelConfig's dataclass defaults are not the trained
architecture — always construct it from config.json as shown above.
Requires torch, safetensors, and python-chess.
Derived from the Lichess January 2022 standard-rated database
(lichess_db_standard_rated_2022-01), following the filtering recipe in the paper's
Appendix C.1:
From 13,813,071 games scanned: 8,329,380 kept for training (≈2.30B tokens) and 20,000 held out for validation. Decontamination against the validation set used 1,483 position keys and removed 79 games.
Elo distribution of the training set peaks in the 1600–2000 range, so play quality reflects typical online club-level chess rather than master play.
python-chess) before using them.The architecture and data recipe follow:
@article{pre2post-chess,
title = {Understanding Reasoning from Pretraining to Post-Training},
year = {2026},
eprint = {2607.16097},
archivePrefix = {arXiv}
}