Downloads · 30 days
0
efe-T/Experiment1-C
Experiment1-C is a machine learning model from efe-T. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
A single continuous pretraining run from random initialization over the whole FineWeb-Edu 10B GPT-2 corpus, in one strictly monotonic pass. The repository contains custom PyTorch source and raw safetensors weights; it…
Downloads · 30 days
0
Access
Public
Updated Sep 10, 2026
Repo size
1.9 GB
Likes
0
Public
Click a slice to open those files.
.py128 KB · 94%
From the Hugging Face model README
A single continuous pretraining run from random initialization over the whole FineWeb-Edu 10B GPT-2 corpus, in one strictly monotonic pass. The repository contains custom PyTorch source and raw safetensors weights; it is not a drop-in Transformers model.
kjj0/finewebedu10B-gpt2 at 440d9d739f970c5d80492e12e81d63d60622ca30, shards
000001 through 000099, read in orderfinewebedu_val_000000.bin, never trained on): not measured1_minus_sqrt cooldown over the final
3,152 steps at the true dataset endwte,
0.3 for the value tables,
0.008 for the head,
0.04 for scalarsEvery run reads checkpoints/latest/manifest.json, restores the newest validated
checkpoint, and continues from the exact sequence it stopped on. There are no
stage boundaries: the optimizer, the learning-rate schedule, and the data cursor
all carry across sessions untouched. A checkpoint is published every
250,000,000 processed tokens.
model/model.safetensors: final standalone model weightscheckpoints/steps/tokens-<n>-step-<s>/: full resumable checkpointscheckpoints/final: the completed runcheckpoints/latest/manifest.json: pointer to the newest valid checkpointcheckpoints/best/manifest.json: pointer to the best-validating checkpointtraining/: exact source and resolved training configurationruntime/: environment metadatametrics/training_metrics.jsonl: per-step training metricsFineWeb-Edu dataset shards are not included.