Downloads · 30 days
0
sup-computer/shakespeare-nanogpt-2
shakespeare-nanogpt-2 is a text generation model from sup-computer. Use it when you need the model to write or continue text. It is set up for nanogpt. The card lists the license as mit.
A sup computer release — a small language model studio. Model page · monorepo (frozen code: projects/shakespeare/models/shakespeare-nanogpt-2/, tag shakespeare-nanogpt-2) · run it in your terminal with the repo's sup…
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2026
Repo size
606 MB
Likes
0
Public
Click a slice to open those files.
.pt359 MB · 59%
From the Hugging Face model README
shakespeare-nanogpt-2 (v2)<div class="takeaways"> <p class="takeaways-label">Key takeaways</p> <ul> <li>The improved model: full corpus + modern architecture (RoPE, RMSNorm, bias-free) + GPT-2 BPE, reaching held-out <code>BPC 1.919</code> (−20% vs. the v1 baseline).</li> <li>The best of Experiment 01's four rounds (Round 3, early-stopped BPE). The Round 4 "champion" that stacked more regularization <strong>regressed</strong>.</li> <li>Still mimicry, and scores are single-seed — <strong>data is the ceiling</strong> the next version will have to raise.</li> </ul> </div>A sup computer release — a small language model studio. Model page · monorepo (frozen code:
projects/shakespeare/models/shakespeare-nanogpt-2/, tagshakespeare-nanogpt-2) · run it in your terminal with the repo'ssupCLI.
The current best model in the shakespeare-nanogpt series:
full corpus + modern architecture + BPE tokenizer, held-out BPC 2.395 → 1.919.
It is the winner of a four-round LLM-assisted research experiment in which
Claude Opus 4.8 acted as the researcher — diagnosing, changing, retraining, and
measuring — while this small model was the thing being improved. This is the
precursor stage to recursive self-improvement,
not RSI itself. v2 is Round 3 of that experiment.
Series note. Successor to
shakespeare-nanogpt-1. Both versions and the full story are inMODELS.md; the write-up is Experiment 01 (report index:research-docs/reports/) and the scoreboard isleaderboard.md.
| Version / git tag | shakespeare-nanogpt-2 |
| Origin | LLM-assisted research experiment, Round 3 (projects/shakespeare/runs/r3-bpe) |
| Architecture | modern — RoPE, RMSNorm, bias-free (core/nanogpt_core/model.py) |
| Size | ~29.9M params |
| Tokenizer | GPT-2 byte-pair encoding (~50k vocab) |
| Checkpoint | models/shakespeare-nanogpt-2/ckpt.pt (weights not committed — rebuild below) |
| Built on | nanoGPT by Andrej Karpathy (MIT) |
| Developed with | Claude Opus 4.8 (Claude Code) as researcher, human oversight |
| License | MIT |
Same as v1 — a learning project, here additionally demonstrating LLM-assisted model development (a large model improving a small one, measured honestly). Generated text is higher-quality Shakespeare-styled mimicry than v1, but still not coherent or factual.
Out of scope: real use of the text; any presentation of output as genuine Shakespeare or as fact. No instruction following, no safety tuning.
The Complete Works of Shakespeare (~5 MB), ~5× the data of v1's Tiny
Shakespeare. Prepared as GPT-2 BPE tokens (projects/shakespeare/data/shakespeare_full_bpe/). A fixed
250k-character held-out test set (projects/shakespeare/test.txt) that no model trains on is
used for evaluation.
Trained with core/nanogpt_core/train.py on the same Apple Silicon MPS setup as v1
(~20 min). Note: the model overfits — validation loss bottomed around step
~1000 then rose; the save-best-val policy automatically kept the early, best
checkpoint.
Scored on the fixed held-out test in bits-per-character (BPC) — a tokenizer-agnostic metric, so char-level and BPE models are directly comparable. Lower is better.
Three levers worked; the fourth backfired. The bars step down through Round 3, then tick back up at Round 4.

| Round | Change | Test BPC | Worked? |
|---|---|---|---|
| — | 1MB control (data-starved baseline) | 2.395 | — |
| 1 | 5× more data (full Complete Works) | 2.036 | yes, −15% |
| 2 | Modern architecture (RoPE + RMSNorm + bias-free) | 2.004 | yes, −1.6% |
| 3 | GPT-2 BPE tokenizer | 1.919 🏆 | yes, −4.3% |
| 4 | "Champion" (+ dropout 0.3 + 4000 iters) | 1.947 | no — regressed |
End to end: BPC 2.395 → 1.919, a 20% reduction. v2 is Round 3, which already combines all three productive levers.
We also tracked Claude's own token cost per round, to ask how much intelligence each unit of improvement cost. Round 1 (fixing the data bottleneck) paid off hugely; later rounds returned far less for similar effort, and Round 4 spent the most and went backwards.

Charts generated by dataviz/.
1337), with no variance estimate. The two smallest table deltas — the
modern-architecture −1.6% and the champion's +1.5% regression — are within
plausible run-to-run noise and should be read as directionally uncertain; the
large gains (more data, BPE) are well outside any plausible noise. The seed is
now a recorded --seed knob, and future versions report multiple seeds.# self-contained v2 folder (weights are gitignored — rebuild them)
cd models/shakespeare-nanogpt-2
python prepare.py # downloads the Complete Works, BPE-encodes it here
python train.py # -> ./ckpt.pt
python eval.py # score on the shared held-out test (expect BPC ~1.919)
python sample.py --start="ROMEO:"
Added in the site-standardization pass (ADR-0015). The card above was unchanged by that pass; this is a tracked addendum. (A later house-style pass, 2026-07-04, edited the card's prose — emphasis and sentence structure only, no facts.) Site-wide fixes — repo links now resolve to GitHub/site routes, code blocks render within the column — apply automatically.