Downloads · 30 days
0
kolbrian/qgate-checkpoints
qgate-checkpoints is a machine learning model from kolbrian. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Training artifacts from an independent research program on query-conditioned attention gating in GPT-2-class transformers. 33 checkpoints, the code that produced them, and the complete result tables — including the ru…
Downloads · 30 days
0
Access
Public
Updated Aug 15, 2026
Repo size
49.4 GB
Likes
0
Public
Click a slice to open those files.
.pt49.4 GB · 100%
From the Hugging Face model README
Training artifacts from an independent research program on query-conditioned attention gating in GPT-2-class transformers. 33 checkpoints, the code that produced them, and the complete result tables — including the runs that did not support the hypothesis.
Code, per-phase result tables, and the full experiment history: github.com/briantkolb/qgate
This repository is an archive first and a model release second. Nothing here is intended for downstream use as a general-purpose language model.
The gate's benefit is a function of how much fixed structure the baseline already has. On a plain GPT-2-class 124M model at 20% warmup it is a large, reproducible improvement. At 1B, on a modern stack, it is marginal and its statistical support is under review. On a short warmup schedule it is worse than no gate at all. On a heavily-optimized speedrun baseline it is decisively harmful. That ordering is the result; the 124M number alone is not.
A second, narrower finding concerns parameterization: the same blend quantity that a 12-parameter per-layer gate learns a clear profile for is one that a 768-parameter per-dimension gate fails to relocate from its initialization at this token budget. Adding capacity made it less learnable, not more.
A single gate applied at the G1 seam — post-SDPA, pre-output-projection:
hs = n_embd // n_head # head_dim
self.qa_gate_proj = nn.Linear(hs * 2, hs, bias=False)
gate = torch.sigmoid(self.qa_gate_proj(torch.cat([q, y], dim=-1)))
y = y * gate # before c_proj (W_O)
At the 124M scale this costs 98,304 parameters (+0.079%). In this corpus the y term is
called A (the attention output); code/model_cheap_qa_minimal.py is a self-contained
single-variant reference implementation with the port surface marked by G1 SEAM banners.
nanoGPT 124M, OpenWebText, 10k iters, warmup 2000, seeds 42/1337/123 (val CE).
Source: results/csv/priority1_runs_20260322_1355.csv.
| variant | added params | val CE | vs baseline |
|---|---|---|---|
| baseline | — | 4.4708 ± 0.0074 | — |
cheap_qa (gate on cat(q, y)) | +98,304 (+0.079%) | 4.3897 ± 0.0051 | −0.0811 (−1.81%) |
full_x (gate on x, the published G1 form) | +7,077,888 (+5.71%) | 4.3950 ± 0.0116 | −0.0758 |
midtier_q (per-head MLP on q) | +196,608 (+0.159%) | 4.4135 ± 0.0086 | −0.0573 |
cheap_q (linear gate on q) | +49,152 (+0.040%) | 4.4198 ± 0.0185 | −0.0510 |
The one comparison here that separates cleanly separates on cost, not loss: cheap_qa
and full_x differ by 0.0053 CE — inside seed noise — and by 72× in added parameters.
Parameter counts are exact; loss deltas are estimates from three seeds.
1. Short warmup — every gate tested loses to the baseline.
Source: results/csv/warmup500_bundle_runs_20260324_1937.csv,
results/csv/phase6b_grok_qafollowups_w500_runs_20260328_0028.csv,
results/csv/phase5_warmup500_toptiers_runs_20260326_1722.csv (warmup 500, 3 seeds each).
| variant | warmup 500 | warmup 2000 | penalty |
|---|---|---|---|
| baseline | 4.5663 | 4.4708 | +0.0955 |
| qa_normed | 4.5902 | 4.3926 | +0.1976 |
| cheap_qa | 4.6130 | 4.3897 | +0.2233 |
| qa_lowrank | 4.6289 | 4.4068 | +0.2221 |
| mlp128_qa | 4.6332 | 4.4032 | +0.2300 |
| full_x | 4.8518 | 4.3950 | +0.4568 |
Short warmup hurts everything, but it hurts every gate roughly twice as much as the baseline, and at warmup 500 the ungated model wins outright. The headline improvement is conditional on a 20% warmup schedule.
2. A hardened baseline — the learned gate harms, decisively.
A 20-run matrix (4 arms × 5 seeds) against the 2025-01-04_SoftCap modded-nanoGPT record on
8×H100, August 2–5 2026. Published noise floor for that record: 3.2791 ± 0.0019 over 80 runs.
| arm | mean val loss | vs baseline |
|---|---|---|
| B — baseline | 3.27790 | — |
| S — static gate (0.5) | 3.27960 | +0.0017 |
| Z — zero-init learned gate | 3.30788 | +0.0300 |
| F — learned gate | 3.30796 | +0.0301 |
F − S = +0.0284, 95% CI [+0.0262, +0.0305], t(4) ≈ 36.5, all five seeds positive, zero distributional overlap (worst ungated 3.2822, best gated 3.3049). Gated arms are also ~7% slower.
The decomposition matters more than the sign: static scaling costs +0.0017; making the
gate learnable costs +0.0284 — 16× more. The damage comes from learning the gate, not
from gating. F − Z = +0.00008 — random and zero init reach the same solution.
This is the best-powered experiment in the corpus, and it is negative. Primary logs
(qgate_run3_MATRIX.tar.gz) are not in this repository; the figures above are from the
August 2–5 session record.
3. Scale — the ordering among gates does not survive.
OLMo-2 1B, OpenWebText, 3 seeds, H100 (rope, rmsnorm, swiglu, QK-norm, post-norm):
| variant | mean ppl | sd |
|---|---|---|
| baseline | 44.137 | 0.179 |
| cheap_qa | 43.703 | 0.124 |
| midtier_q | 43.683 | 0.159 |
cheap_qa and midtier_q differ by 0.02 ppl — p = 0.87, a tie — despite clear separation at
124M. See the caveats below before using the 1B result for anything.
Two or more files here disagree. These are shown rather than resolved.
The A10 comparison. Two boards, same hardware, same recipe, different conclusions:
| source | baseline | cheap_qa | midtier_q | ordering |
|---|---|---|---|---|
results/summaries/priority1_summary_20260322_1355.txt (A10 ref block) | 4.4675 | 4.4151 | 4.4039 | midtier_q wins |
results/phase8_warmup2000_bulk_a10_runs_20260329_0650.csv (19 variants, 57 runs) | 4.4676 | 4.4155 | 4.4183 | cheap_qa wins |
Baseline and cheap_qa agree to 0.0004 across both. midtier_q differs by 0.0144, and that
single number is what the "ordering inverts on A10" claim rests on. The larger and later
board does not reproduce the inversion. Additionally, the A10 lane ran a different software
stack (torch 2.7.0+cu128) from the canonical 3060 board (2.5.1+cu121), and the project record
explicitly blocks a hardware-only interpretation until a version-matched rerun exists.
What is consistent across both A10 boards: cheap_qa beats baseline, and it drops from
1st on the 3060 to 7th of 19 on the A10 (behind x_full_headspec at 4.3858).
The 1B result. An August 4 review found: the gate-norm "stability" claim retracted
(post-training ‖W‖_F is statistically indistinguishable from an untouched init draw,
P = 0.494); the win record mislabelled (cheap_qa alone is 3/3, sign-test p = 0.125 — not
significant; the gated family pooled is 6/6, p = 0.0156); seed 42 resumed from a
step1536 checkpoint in all three arms and carries the largest margin (dropping it moves the
mean delta 0.433 → 0.285); the advantage is late-emerging (baseline leads on 2/3 seeds at
step 500); midtier_q logged no gate proxy, so there is no cross-variant control. Whether
the OLMo gate trained at all is not established — a frozen random projection is not
excluded. The checkpoints were lost with the rented instance, so this may stay open.
Whether the 1B run left LR warmup. One record says the schedule was truncated but
t_warmup fixed 200M→40M with expected_max_steps: 1526; another says warmup was never
exited. The rendered per-run config was never recovered. Unresolved.
The WSL lane. Two docs report baseline/cheap_qa as 4.4759/4.3902 and 4.4732/4.3800. The WSL baseline also shows late-training spikes attributed to WSL2 memory management — contamination in the direction that widens the gap. This lane also ran torch 2.7.0.
Seed sensitivity. On out-of-band seeds 0/1/2 (results/phase9_...20260328_2327.csv),
cheap_qa is 4.4139 ± 0.0094 rather than 4.3897, and qa_normed (4.4103) beats it. The
±0.0051 above is a within-seed-set figure for 42/1337/123.
Full vectors, all 12 layers, 3 seeds — 4,608 values per condition. Note that per-layer-mean bands are much tighter than the individual-value bands and the two are easy to conflate:
| phase | init | individual values | per-layer means | mean | val CE |
|---|---|---|---|---|---|
| 19b | 0.50 | — | q 0.4587–0.5202 / y 0.4687–0.5766 | 0.481 / 0.504 | 4.4015 |
| 19d free | ~0.50 | 0.3593–0.6821 | 0.4558–0.5718 | 0.4924 | 4.4085 |
| 19c informed | 0.74 | 0.5745–0.8457 | 0.6995–0.7826 | 0.7317 | 4.3950 |
What holds: the mean does not move from its initialization — 0.74 → 0.7317, 0.50 → 0.4924. Initialization sets where the distribution sits.
What does not hold: "the scales do not move." They spread substantially — 19c covers a 0.27 range, 19d covers 0.32. Individual dimensions differentiate; the center does not shift.
The mechanism claimed earlier — a flat loss direction — is contradicted by this repository. The same blend quantity moves decisively under coarser parameterization:
| parameterization | params | init | converged |
|---|---|---|---|
| phase 14, single scalar | 1 | 0.5 | 0.7391–0.7443 |
| phase 17, per-layer | 12 | 0.5 | q 0.3803–0.5551 (mean 0.4481) / y 0.4126–0.6529 (mean 0.5144) |
| phase 19b/c/d, per-dimension | 768 | 0.50 / 0.74 | mean stays at init |
A scalar moves +0.24 from its initialization. Twelve per-layer values differentiate across a 0.27 range and reproduce a consistent profile. Seven hundred sixty-eight per-dimension values do not relocate their mean. The loss is not flat along this axis — the fine parameterization is not identified at this budget. Gradient dilution across 1,536 parameters and simple undertraining are both live explanations and this corpus cannot separate them.
cheap_qa's margin is seed-set dependent — 4.3897 on seeds 42/1337/123, 4.4139 on
seeds 0/1/2.cheap_qa seeds, and 0 of 3 baseline seeds, ended with their best
validation loss at the final evaluation. Most runs were drifting upward late.⚠️ results/phase19b_canonical_*_0709.* is a failed run — six seeds, returncode=2,
best=nan, no checkpoint. It sits beside the real six-seed data (..._0711).
⚠️ results/summaries/p2_rerun_summary_20260323_1714.txt ranks full_x first at 4.3870 on a
two-seed partial. results/summaries/full_x_seed123_result_20260324_1712.txt (4.3950,
3 seeds) supersedes it.
⚠️ results/summaries/a10_replication_summary_20260322_0639.txt ends with an auto-generated
PAPER STATEMENTS block asserting hardware independence and |delta| < 0.02. Its own table
forty lines above shows deltas to −0.1257 and prints Hardware consistency: INVESTIGATE ✗.
Trust the tables, not the prose blocks.
⚠️ Era warning. Results predating the beta2=0.99 fix land in the 4.5–4.6 CE band and are
not comparable to clean-recipe results in the 4.38–4.47 band.
⚠️ checkpoints/canon/out-shakespeare-char/ is the upstream nanoGPT demo, not part of this
study.
Gating the SDPA output at G1 is established. This is a replication and mechanism study, not a novelty claim.
full_x here is that formulation ported to nanoGPT.H is the y/A term in
cat(q, y). Subsequent and independent; their 5B/500B result reaches a regime this corpus
cannot.Two of Qiu et al.'s qualitative findings reproduce here at ~1/10,000 of their token budget:
G1 gating beats baseline (6.026 → 5.761 PPL there; 4.4708 → 4.3897 CE here, at warmup 2000),
and input-dependent beats input-independent (5.917 → 5.761; 4.4483 → 4.3897). The
input-independent control is an exact structural match — a zero-initialized learnable
(n_head × head_dim) parameter through a sigmoid in both cases.
cheap_qa was confirmed 2026-03-20
(results/csv/cheap_qa_confirmation_runs_20260320_1854.csv), before the prior-art search that
located Qiu et al.; the project was then reclassified from a novelty claim to a replication
study. This establishes no priority — these results were private until August 2026.
checkpoints/
canon/ 21 x ckpt.pt base run + phases 14-19d (RTX 3060, Windows)
wsl_bridge/ 12 x ckpt.pt phases 10-12 (RTX 3060, WSL)
code/ model/train for both lanes + the minimal champion module
results/ phase 7-19d run CSVs, summaries and logs
SHA256SUMS.txt integrity manifest for all 33 checkpoints
nanoGPT-format ckpt.pt at fp32, ~1.4 GiB each, optimizer state included.
sha256sum -c SHA256SUMS.txt
cheap_qa fall from 1st to 7th on the A10 while still beating baseline?Derived from nanoGPT by Andrej Karpathy (MIT).
See LICENSE.
2026-08-15 — Card rewritten to cover the full corpus: the August speedrun matrix, the warmup sweep, and the per-dimension bands computed from the full vectors. The first version described only the 124M and 1B lanes, and stated the improvement without its warmup and baseline-hardness conditions.