Downloads · 30 days
54
51% of all-time downloads
PursuitOfDataScience/argonne-3.5-base
argonne-3.5-base is a text generation model from PursuitOfDataScience. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Argonne 3.5-base is a 2.88B-parameter decoder-only transformer trained from scratch. It is the base (foundation) checkpoint of the Argonne 3.5 line and the successor to argonne-3.0-base.
Downloads · 30 days
54
51% of all-time downloads
All-time downloads
106
Public
Parameters
2.9B
5.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.8 GB · 100%
From the Hugging Face model README
Argonne 3.5-base is a 2.88B-parameter decoder-only transformer trained from scratch. It is the base (foundation) checkpoint of the Argonne 3.5 line and the successor to argonne-3.0-base.
Two things separate it from 3.0-base:
The architecture is unchanged from 3.0 — grouped-query attention with QK-norm, V-norm, sandwich norms, interleaved local/global attention, and a final logit softcap. What changed is the training recipe (FP8, a higher peak LR with a proper cooldown, tighter gradient clipping) and the data curriculum.
This is a base model: no instruction tuning, no alignment, no safety filtering.
Looking for the reasoning model? This base was post-trained into Argonne-3.5-think, which scores 65.00 / 73.00 greedy and 74.00 / 82.67 self-consistency on clean SVAMP / ASDiv — vs 22.00 / 32.33 and 30.00 / 39.33 for the previous generation, measured head-to-head in one job.
| Component | Specification |
|---|---|
| Parameters | 2,882,162,688 (~2.88B) |
| Layers | 24 transformer blocks |
| Hidden size | 3,072 |
| Attention heads | 12 query / 4 key-value (GQA) |
| Head dimension | 256 |
| Feed-forward | SwiGLU MLP, 8,192 intermediate dim |
| Attention pattern | Interleaved local/global causal attention |
| Local attention window | 256 tokens (every other layer) |
| Normalization | RMSNorm with QK / V / sandwich norms |
| Position encoding | RoPE (θ = 1,000,000) |
| Logit stabilization | Final logit softcap = 15.0 |
| Context length | 13,568 tokens (trained, not extrapolated) |
| Vocabulary size | 151,669 |
| Tied embeddings | Yes (input ↔ output) |
Three stages, all causal language modeling, all on 3× NVIDIA H100/H200 GPUs with DDP.
| Stage 1 — pretrain | Stage 2 — reasoning anneal | Stage 3 — context extension | |
|---|---|---|---|
| Script | pretrain.py | continue_pretrain.py | continue_pretrain.py |
| Steps | 50 → 244,000 | 244,010 → 308,730 | 308,740 → 321,062 |
| Tokens | 65.30B | 17.50B | 6.02B |
| Cumulative | 65.31B | 82.81B | 88.84B |
| Sequence length | 1,024 | 1,024 | 13,568 |
| Batch / GPU | 38 → 44 | 88 | 3 |
| Grad accumulation | 2 | 1 | 4 |
| Effective batch | 233,472 → 270,336 tok/step | 270,336 tok/step | 488,448 tok/step |
| Peak LR | 6.0e-4 | 2.0e-4 | 1.0e-4 |
| End LR | 6.0e-5 | 2.0e-5 | 1.0e-5 |
| Warmup | 8,000 steps | 0 | 0 |
| Schedule | WSD, cooldown_frac 0.15 | cooldown to 0.1× | cooldown to 0.1× |
Shared across all three stages:
| Item | Value |
|---|---|
| Optimizer | AdamW (β₁=0.9, β₂=0.95, weight decay 0.1) |
| Gradient clipping | 0.4 |
| Precision | FP8 (torchao tensorwise, including lm_head) under bf16 autocast; fp32 optimizer states |
| Vocab padding | 151,669 → 151,680 during training for the FP8 lm_head (trimmed back on export) |
torch.compile | Enabled |
| Gradient checkpointing | Enabled |
| Data parallel | 3 GPUs (DDP) |
| Total optimizer steps | 321,062 |
| Final train loss | 0.8923 (stage-3 slice average; not comparable across stages — the mixtures differ) |
| Checkpoint dtype on Hub | bfloat16 |
| Weight format on Hub | 5 sharded safetensors + index |
3.0-base ran WSD with cooldown = 0 — the stable phase only, no decay. 3.5 uses a real cooldown
in every stage (visible in the LR panel of the figure below), a higher peak LR (6e-4 vs 3e-4), a
longer warmup (8,000 vs 1,000 steps), and tighter gradient clipping (0.4 vs 1.0). The tighter clip
and QK-norm are what make the higher LR stable.
| Stage | Corpus | Tokens |
|---|---|---|
| 1 — pretrain | FineWeb + FineMath, 85/15 | 65.30B |
| 2 — reasoning anneal | code / math / reasoning / tool mixture with a general-web replay tier (below) | 17.50B |
| 3 — context extension | a disjoint slice of the same stage-2 composite, read at 13,568 tokens | 6.02B |
The stage-2/3 composite is built by
build_reasoning_corpus.py
from:
| Tier | Source |
|---|---|
| code | nick007x/github-code-2025 (≥2 stars) · nvidia/Nemotron-Competitive-Programming-v1 |
| math | nvidia/OpenMathReasoning |
| reasoning | a-m-team/AM-DeepSeek-R1-Distilled-1.4M · open-r1/Mixture-of-Thoughts · PursuitOfDataScience/0.5M-thinking |
| tool | nvidia/Nemotron-SFT-Agentic-v2 |
| general (replay) | HuggingFaceFW/fineweb-edu |
Reasoning traces keep their <think>/tool tags, so the base has seen that formatting before any
fine-tuning. The corpus is decontaminated against common evaluation sets, and stages 2 and 3 use
disjoint slices of it (--holdout_frac / --part) so stage 3 is not a second epoch over
stage 2's tokens.
The general-web replay tier exists because it is needed: an earlier build capped that tier at 2.2B tokens (9% of the anneal mix), and 18% of the way in, held-out FineWeb-Edu cross-entropy had risen +0.246 nats — the base was measurably regressing on general text while improving on the target tiers. The tier was raised to a 20.6B-token pool.
Tokenizer: Qwen/Qwen3-0.6B-Base (151,669-token
vocab), via the Qwen2Tokenizer compatibility class. Bundled with the checkpoint.

Loss, perplexity, and learning rate against cumulative tokens across all three stages, with the stage boundaries marked. The LR panel shows the three cooldowns. Note that the loss step down at each stage boundary is a change of data mixture, not a capability jump — the anneal and context-extension corpora are intrinsically lower-entropy than FineWeb, so cross-stage loss values are not comparable.
Two measurements were run on the final checkpoint. Both are reported with their limitations, because neither is a general capability benchmark.
The question is whether stage 3 actually taught the model to use long positions, or just trained it more. The test is a paired A/B on identical held-out windows (40 documents, 24,576 tokens each, proof-pile-2 arXiv — a domain neither stage trained on), comparing the final weights against the stage-2 checkpoint they were seeded from. Lower is better.
| Token position | Stage-2 checkpoint (ctx 1,024) | Argonne 3.5-base (ctx 13,568) |
|---|---|---|
| 0 – 1,024 | 2.194 | 2.161 |
| 1,024 – 2,048 | 5.536 | 1.860 |
| 2,048 – 4,096 | 5.969 | 1.576 |
| 4,096 – 8,192 | 5.895 | 1.320 |
| 8,192 – 13,568 | 5.961 | 1.207 |
| 13,568 – 20,480 | 5.965 | 1.122 |
| 20,480 – 24,576 | 5.938 | 1.096 |
Three things this shows:
Reproduce with
reasoning/exp_longctx_learning.py.
A 35-item greedy few-shot probe (20 arithmetic/word-problem, 15 world-knowledge) used as a go/no-go gate for whether a base is worth running a reasoning recipe on.
| Checkpoint | Math /20 | General /15 |
|---|---|---|
| Stage-2 (pre-extension, step 308,733) | 18 | 15 |
| step 320,885 | 18 | 15 |
| step 321,054 | 17 | 15 |
| step 321,062 (this model) | 18 | 15 |
Both axes clear the ≥14/20 ∧ ≥14/15 gate, and the context-extension stage cost nothing on either.
Read this as a gate, not as a capability number. The probe is small (n=20/15), it saturates,
and it has a measured ±2-item noise floor — which is why three checkpoints are shown rather than
one. It says the base is worth building on. It does not say how good it is. Reproduce with
reasoning/probe_pretrain_ckpt.py.
No standard held-out benchmark suite (MMLU, ARC, HellaSwag, GSM8K, …) has been run on this checkpoint yet. Those numbers are not withheld — they do not exist, and this card will be updated when they do. Do not infer benchmark standing from the two measurements above.
One caution specific to this line: GSM8K is contaminated for Argonne reasoning derivatives downstream of this base, and should not be used to grade them.
Built from the GitHub main branch: https://github.com/PursuitOfDataScience/ArgonneAI/tree/main
| File | Role |
|---|---|
model.py | ArgonneModel / ArgonneConfig architecture + KV cache (bundled here as model.py) |
pretrain.py | stage 1 — DDP pretraining loop |
continue_pretrain.py | stages 2 and 3 — anneal and context extension |
build_reasoning_corpus.py | builds the stage-2/3 corpus (tiering, decontamination, disjoint slicing) |
reasoning/probe_pretrain_ckpt.py | the two-axis base gate probe |
reasoning/exp_longctx_learning.py | the position-bucketed long-context NLL probe |
reasoning/thinking_training.md | the full lab notebook for the reasoning line |
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "PursuitOfDataScience/argonne-3.5-base"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
)
prompt = "Write a short paragraph about scientific computing at Argonne National Laboratory."
inputs = tokenizer(prompt, return_tensors="pt")
input_ids = inputs["input_ids"].to(model.device)
output_ids = model.generate(
input_ids,
max_length=input_ids.shape[1] + 128,
temperature=0.8,
top_p=0.95,
top_k=50,
do_sample=True,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
trust_remote_code=True so the custom ArgonneModel / ArgonneConfig classes
(model.py) are registered. Unlike the 3.0-base card, this repo ships an auto_map in
config.json, so from_pretrained resolves the classes without any manual setup.generate method on ArgonneModel takes max_length (total sequence length), not
max_new_tokens.model.safetensors.index.json weight map.lm_head.weight is reported missing on load. This is expected and benign — embeddings are tied
(tie_word_embeddings: true), so lm_head takes its weights from embed_tokens..generate().do_sample=False) for deterministic output.@misc{argonne35base,
author = {PursuitOfDataScience},
title = {Argonne 3.5-base},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/PursuitOfDataScience/argonne-3.5-base}
}