Downloads · 30 days
6
18% of all-time downloads
avewright/chess-transformer-200m-maxelo
chess-transformer-200m-maxelo is a other model from avewright. Use it for the other task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Finetune of avewright/chess-transformer-200m-compact-soft for high Elo without MCTS: next-move prediction via legal-mask policy argmax.
Downloads · 30 days
6
18% of all-time downloads
All-time downloads
34
Public
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
.pt1.6 GB · 100%
From the Hugging Face model README
Finetune of avewright/chess-transformer-200m-compact-soft
for high Elo without MCTS: next-move prediction via legal-mask policy argmax.
| Architecture | ChessTransformer ~204M (fused board encoder, 16L × 1024d × 16H) |
| Vocab | compact (MOVE_VOCAB_VERSION=compact) |
| Heads | Spatial policy + 3-class WDL value |
| Inference | Policy argmax only (no search) |
| Base | chess-transformer-200m-compact-soft |
| Train run | exp189 (soft MultiPV + deep soft mix + hard depth≥15) |
| File | Meaning |
|---|---|
best_model.pt | Best blended soft holdout top-1 during deep-mix resume |
latest_model.pt | Shutdown weights at step 2851 |
config.json | Architecture + training metadata |
PROGRESS.md | Full session write-up |
elo_eval.json | Raw Elo ladder results |
Lean checkpoints contain model_state_dict + metadata (no optimizer).
avewright/exp186-sf-multipv-2mavewright/exp190-phase-deep-soft (~40% of soft steps)min_depth ≥ 15Augmentation: horizontal flip on soft batches (hflip_p=0.5).
Evaluated with elo_eval_latest.py vs Stockfish 18 UCI_LimitStrength
(50ms/move, opening book + Syzygy, 8 openings × both colors):
| Opponent Elo | Score | W–D–L |
|---|---|---|
| 1500 | 0.625 | 7–6–3 |
| 1800 | 0.438 | 4–6–6 |
Estimated Elo ≈ 1700 (bracket 1500–1800; small sample, noisy).
import os, torch
os.environ["MOVE_VOCAB_VERSION"] = "compact"
from play import load_model # or elo_eval_latest.load_eval_model
model = load_model("best_model.pt", device="cuda")
model.eval()
Or play in the local GUI:
export MOVE_VOCAB_VERSION=compact
python play_factory_gui.py --checkpoint best_model.pt
Training code: avewright/transform —
experiments/exp189_200m_maxelo_policy.py, docs/PROGRESS_2026-07-10.md.