Downloads Ā· 30 days
0
jbduran/bartholomew-experiments
bartholomew-experiments is a machine learning model from jbduran. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
The complete training archive behind BART ā 39 runs, their checkpoints, tokenizers, and evaluations, including every dead end.
Downloads Ā· 30 days
0
Access
Public
Updated Aug 21, 2026
Repo size
1.2 TB
Likes
1
Public
Click a slice to open those files.
.pt1.3 TB Ā· 100%
From the Hugging Face model README
The complete training archive behind BART ā 39 runs, their checkpoints, tokenizers, and evaluations, including every dead end.
š Read the write-up Ā· š Unbounded Labs
If you want the model itself, use bart or bart-sft. This repository is for reproducing or inspecting how they were reached.
experiments/ 39 runs ā each owns its tokenizer and base checkpoints
evaluations/ vintage-core results across models
archive/ legacy paths (pre-lineage-v1)
Each base experiment owns its tokenizer and base checkpoints. SFT runs nest under their exact base
parent (experiments/<base>/sft/<sft-run>/), and post-training runs nest under their exact SFT
parent.
| Fragment | Meaning |
|---|---|
d12, d24, d32 | model depth |
r11ār30 | target parameter-to-data ratio |
ctx4096, ctx8192 | max sequence length |
sssl | window pattern |
fulltok, randtok | tokenizer variant |
The prefix tells you the corpus generation:
| Prefix | Corpus |
|---|---|
think- | bart-dataset-v1 |
thinkcleaned- | v2 |
clean1930s- | v3 |
Think.Unbounded- | v3 plus midtraining |
Think.Unbounded-d32 is the first d32 run, trained on midtrain mixtures whose ratios were computed
by document count. Those blends under-delivered badly ā one targeting 60% midtrain supplied
25%, one targeting 30% supplied 11.5%.
Think.Unbounded-d32-v2mix-cont is the repair: it branches from that run at step 5500, carries the
optimizer state, and continues on corrected token-based mixtures. It is the model shipped as
bart. The two configs side by side are the clearest record
of that bug and its fix.
Its sft/ directory holds six fine-tuning variants. The one shipped as
bart-sft is
pre1930-curriculum-c3-robust-v2; the others ā karpathy-modern-sft-v1,
nanochat-default-datamatch-v1, c3-robust, and two c3-robust-v3 configs ā are kept here.
clean1930s-d24-r12-ctx4096-sssl-fulltok-v1 is the d24 / 1.38B run; its tokenizer is the one every
midtrain mixture was built against.
Note. Each
config.jsonrecords repo and dataset names as they were at training time (jbduran/think.nano,think-dataset-clean-1930s,think-midtrain). Repo renames still resolve via Hub redirects, but the midtrain mixture subfolders they reference (mixed/v2/ratio_21/data) moved tomixtures/v2-by-tokens/ratio_21/datawhen that dataset was restructured. Configs are kept as historical records rather than rewritten.
bart Ā· bart-sft Ā· bart-dataset-v3 Ā· bart-midtrain Ā· bart-dataset-scripts Ā· bart-midtrain-scripts
MIT
Built by Unbounded Labs.