Downloads · 30 days
0
AdaptiveChunking/hnet-chunking-spec-82m
hnet-chunking-spec-82m is a machine learning model from AdaptiveChunking. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Thirty-one byte-level H-Net runs at 82,568,832 parameters, one hierarchy stage, trained to study what a learned chunker discovers when it is free to choose its own computational granularity — and, unlike the 22.5M pil…
Downloads · 30 days
0
Access
Public
Updated Sep 12, 2026
Repo size
246 GB
Likes
0
Public
Click a slice to open those files.
.safetensors254 GB · 100%
From the Hugging Face model README
Thirty-one byte-level H-Net runs at 82,568,832 parameters, one hierarchy stage, trained to study what a learned chunker discovers when it is free to choose its own computational granularity — and, unlike the 22.5M pilots, when it discovers it.
The pilots at AdaptiveChunking/hnet-chunking-pilots
retained only a final checkpoint. These runs retain the whole log-spaced
trajectory, which is the point: it is what makes training-dynamics and
activation-patching questions answerable at multiple steps rather than at one.
That trajectory cannot be reconstructed after the fact.
| run | condition | seed | steps | training bytes | checkpoints |
|---|---|---|---|---|---|
spec_A_s{0,1,2} | A — baseline: global ratio loss throughout, router never frozen | 0,1,2 | 11,000 | 0.72 GB | 21 each |
spec_B_s{0,1,2} | B — the schedule: global → parity-A → freeze | 0,1,2 | 11,000 | 0.72 GB | 21 each |
spec_C_s{0,1,2} | C — reversion control: global → parity-A → global, router left unfrozen | 0,1,2 | 11,000 | 0.72 GB | 21 each |
spec_D_s{0,1,2} | D — parity-A in all three phases, no freeze | 0,1,2 | 11,000 | 0.72 GB | 21 each |
spec_E_s{0,1,2} | E — global → parity-B (max-min) → freeze | 0,1,2 | 11,000 | 0.72 GB | 21 each |
specfull_A_s{0,1,2} | A, at the full data budget | 0,1,2 | 76,300 | 5.00 GB | 30 (s0) / 44 (s1, s2: extra probe every 500 steps to 8k) |
specfull_B_s{0,1,2} | B, at the full data budget | 0,1,2 | 76,300 | 5.00 GB | 30 / 44 / 44 |
exp32_onechunk_s{0,1,2} | A, trained with one chunk per 1024-byte sequence (main net reduced to a per-sequence vector) | 0,1,2 | 11,000 | 0.72 GB | 21 each |
exp32_local_s{0,1,2} | A, sliding-window (W=128) attention in the byte encoder/decoder | 0,1,2 | 11,000 | 0.72 GB | 21 each |
exp32_local_onechunk_s{0,1,2} | both of the above | 0,1,2 | 11,000 | 0.72 GB | 21 each |
exp32_onechunk_full_s0 | one chunk per sequence, at the full data budget | 0 | 76,300 | 5.00 GB | 30 |
The exp32_* runs are controls (experiments/exp32_train_controls.py monkey-patches the forward;
weights load into the same HNet class but must be run with the same patch to reproduce their
BPB — see results/exp32_trained_controls/RESULTS.md).
Data, seed schedule and FLOP budget are held identical within a scale; only the
chunker's objective varies. Architecture: d_enc 384, d_main 768, n_main 10,
seq_len 1024, batch 64 (specfull), corpus FineWeb2 + FineWeb-v1 (en).
The parity effect replicates at 3.7× the pilot's parameters, and at the full data budget, with essentially no likelihood cost:
| scale | Gini(chunks/sentence) A → B | reduction | high-resource BPB cost |
|---|---|---|---|
| pilot, 22.5M, 0.2 GB, 3 seeds | 0.2195 → 0.0831 | −62.1% | +0.38% |
| spec, 82.5M, 0.72 GB, 3 seeds | 0.2181 → 0.0997 | −54.3% | −0.82% (B is better) |
| specfull, 82.5M, 5.00 GB, 3 seeds | 0.2298 → 0.1886 ± 0.073 (per seed 0.104 / 0.181 / 0.281) | −17.9% | +0.22% |
At 0.72 GB: same direction, same magnitude, same chunks-per-sentence vs chunks-per-character asymmetry as the pilot, and condition C's lock-in reversion (+0.115) replicates to the decimal. At 5 GB the freeze does not bank the effect: the router's W_q/W_k are frozen at step 53,410 but the encoder that feeds them is still trained, and in two of three seeds the allocation drifts back toward the unregularised one over the 22,890 frozen steps (end-of-phase-2 Gini 0.066–0.103 in all three; end-of-run 0.104 / 0.181 / 0.281). The equalisation is a property of the objective being on, not a state that can be frozen in. Mean BPB drops from 1.376 (0.72 GB) to 1.153 (5.00 GB), so the full-budget models are genuinely better trained, not just longer.
Each run directory holds {step:06d}.safetensors (one per retained checkpoint),
config.json, final.json and log.jsonl. config.json carries model_config,
the training config, a checkpoints list mapping tag → step, and a
tied_weights map. emb.weight and head.weight are tied, so only emb.weight
is stored — restore with:
import json
from safetensors.torch import load_file
sd = load_file("spec_A_s0/011000.safetensors")
cfg = json.load(open("spec_A_s0/config.json"))
for dst, src in cfg["tied_weights"].items():
sd[dst] = sd[src]
manifest.json at the repo root lists every run with its parameter count,
condition, seed, final step and checkpoint count.
Probe trajectories for these runs (253 npz per 11k-step run, 361 / 529 per full run)
live alongside the pilots' in
AdaptiveChunking/hnet-chunking-probes
under the same spec_* / specfull_* names.
These are documented rather than hidden, because they change what the logged fields mean:
mask and mask_level1 are byte-for-byte identical. Everything here is a
single-stage model (S=1); mask_level1 is a placeholder, not a second
hierarchy level. Any analysis treating it as one is measuring the same thing
twice.len(mask) runs 21 bytes long relative to the text it indexes.freeze_scope: router+encoder — the arm the 5 GB freeze failure calls for —
has not been run.Produced by tier1/train.py in beetle-hnet;
launched by tier1/launch_spec_overnight.sh, launch_spec_cde.sh,
launch_spec_full.sh, launch_spec_full_seeds.sh and (exp32)
scratch/logs/r4/lane_exp32.sh; packaged by tier1/push_spec_to_hf.py. Analysis
in results/exp21_tier1_spec/ and results/exp32_trained_controls/.