Downloads · 30 days
0
robbiethompson2018/nd-rl-checkpoints
nd-rl-checkpoints is a reinforcement learning model from robbiethompson2018. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product.
Experimental small proof models from the nd-rl pretraining autoresearch workstream. This repository shares checkpoint weights and their recorded experiment summaries.
Downloads · 30 days
0
Access
Public
Updated Sep 30, 2026
Repo size
92.9 GB
Likes
0
Public
Click a slice to open those files.
.pt92.7 GB · 98%
From the Hugging Face model README
Experimental small proof models from the nd-rl pretraining autoresearch workstream. This repository shares checkpoint weights and their recorded experiment summaries.
runs/<run>/pretrain.pt: immediately after pretraining, before RL.runs/<run>/rl_round_XX.pt: after the named RL round.runs/<run>/final.pt: after pretraining and reinforcement learning, with final evaluation
completed. These are not pretraining-only weights.From run 012 on, each seed run is one directory, runs/<experiment>_s<seed>; where present,
runs/<experiment> holds only the merged multi-seed summary. Earlier runs are single-seed
directories. The autoresearch loop uploads every seed run when it finishes.
previous_attempts/ keeps an archived, cancelled attempt whose rerun is the scored one, and
pretrain_poc/ the pretraining proof of concept. The earliest
runs predate stage checkpoints and have only final.pt, or no weights at all. Each
summary.json records the run's status, source hashes, budgets, and evaluation results. Scores
across different harness hashes or budgets should not be treated as directly comparable.
hf download robbiethompson2018/nd-rl-checkpoints --include 'runs/001_gpu-batching/*' --local-dir checkpoints
import torch
checkpoint = torch.load(
'checkpoints/runs/001_gpu-batching/final.pt',
map_location='cpu',
weights_only=True,
)
# Construct the matching model architecture, then:
# model.load_state_dict(checkpoint['state'])
# checkpoint['tok_mode'] specifies the tokenizer mode.
These are raw PyTorch state dictionaries, not Transformers from_pretrained packages.
They do not contain optimizer or RNG state for exact training resumption.
Checkpoints do not embed a full architecture configuration. Most run directories include
source/ with the exact pretrain.py, harness.py, and remote.sh that produced them. For
the earliest runs, the top-level sources/ folder contains available pretraining
implementations; match the first 12 hexadecimal characters of a file's SHA-256 to
pretrain_sha in the run summary. baseline.py matches run 001_gpu-batching,
010_replicate-seed1, and the s20_baseline_* runs. Do not assume every run uses the same
architecture. The external training fork is dan-pandori/nd-takehome, pinned in each run's
source/remote.sh (FORK_COMMIT): ab629c7d2cdf3dc50d471e302d14643ba0782749 for runs before
098, and 51604b3feaa5e66821d1faba22238fe6d96d46d0 (Lean 4 tokenizer and reward) from 098 on.
The autoresearch metric uses development theorems and is not a final held-out benchmark. No inference service or hosted GPU is provisioned by this repository.