Downloads · 30 days
0
shazmate/gpct-v2
gpct-v2 is a machine learning model from shazmate. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A static-site harness for testing a chess AI that plays by predicting the next token of a chess game. The AI isn't here yet — a mock model stands in behind the exact interface the real one will use. No build step, no…
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
51.5 GB
Likes
0
Public
Click a slice to open those files.
.pt1.1 GB · 100%
From the Hugging Face model README
A static-site harness for testing a chess AI that plays by predicting the next token of a chess game. The AI isn't here yet — a mock model stands in behind the exact interface the real one will use. No build step, no dependencies, hostable on GitHub Pages as-is.
node tools/serve.mjs # -> http://localhost:4173
(any static server works; ES modules just can't load from file://)
GitHub Pages: push, then Settings → Pages → Deploy from a branch →
main / root. The site is fully static.
quality set — the model is now conditioned on "this move is an
inaccuracy / mistake / blunder". Nerf tokens are masked out on the retry.?! / ? / ??
glyph in the move list if it was nerfed.The Model output panel in the UI shows the sampled chain and the top of the masked distribution at every step of the last computer move.
A move is a sequence of modifier tokens followed by one core move token — the core token is what ends a move:
[#|+]? [x]? [=Q|=R|=B|=N]? [core]
fxe8=Q+ -> [+] [x] [=Q] [fe8]
Nxf3 -> [x] [Nf3]
Raxe1# -> [#] [x] [Rae1]
e4 -> [e4]
O-O -> [O-O]
Cores keep SAN's disambiguation so nothing is lossy: fe8 is a pawn capture
core (source file + destination), piece cores may carry file / rank / square
disambiguators (Nbd7, R1e2, Qh4e1). sanToTokens() / tokensToSan()
in js/vocab.js convert both ways.
The vocabulary is 5,252 tokens:
| group | count |
|---|---|
| core moves (pawn 176, N/B/R/Q/K 5,064, castling 2, generated geometrically) | 5,242 |
modifiers + # x =Q =R =B =N | 7 |
nerf <inaccuracy> <mistake> <blunder> | 3 |
Core generation is a deterministic, slightly loose superset of real SAN
(every core is geometrically possible on an empty board; impossible ones are
just permanently masked). tools/vocab-check.mjs plays random games and
asserts every legal move tokenizes into the vocabulary and round-trips:
node tools/vocab-check.mjs 500
A model is any object with an async predict(ctx) returning either a
Float32Array / number[] of length VOCAB_SIZE (index i = probability
of TOKENS[i]) or a plain {token: probability} map. It is called once
per token:
{
fen, // current position
moves, // SAN history: ['e4', 'c5', ...]
historyTokens, // same history as a flat token stream, nerf tokens included
moveTokens, // tokens emitted so far for the move being decoded
turn, // 'w' | 'b' — the side the model plays
opponent, // { rating: 1500 | null, site: 'lichess' | null }
legalMoves, // SAN list, provided for convenience — free to ignore
quality, // null, or 'inaccuracy'|'mistake'|'blunder' after a nerf
vocab, // the token list
}
Outputs don't need to be pre-masked or normalized — the engine masks illegal tokens to zero and renormalizes whatever you return.
Swap the model by replacing js/mock-model.js (imported in js/app.js), or at runtime — useful while weights load asynchronously:
window.chessGpt.setModel({ name: 'the real one', async predict(ctx) { ... } });
Other seams: window.chessGpt.loadFen(fen) jumps to a position,
window.chessGpt.config exposes ENGINE_CONFIG (sampling temperature,
disable nerf tokens, think-delay).
index.html page shell
style.css minimalist theme
js/vocab.js token vocabulary + sanToTokens / tokensToSan <- swap for your tokenizer
js/engine.js mask -> renormalize -> sample decode loop <- the AI's socket
js/mock-model.js placeholder model <- replace with the real AI
js/app.js game flow, UI wiring
js/board.js SVG board rendering + input
js/pieces.js geometric piece shapes
lib/chess.js vendored chess.js 1.4.0 (rules, legality, SAN)
tools/serve.mjs dev server: node tools/serve.mjs
tools/vocab-check.mjs vocabulary/round-trip validation
TRAINING.md is the plan for building the training set from evaluated Lichess games (eval-derived nerf labels, this vocabulary, nanoGPT packing).
900 lichess
blunders a lot; leave rating empty for a middling default). The rating
site is plumbed through but unused by the mock — a real model might embed
it, since ratings aren't comparable across sites.ENGINE_CONFIG.temperature, default 1.0).