Downloads · 30 days
54
23% of all-time downloads
Ace-2504/slm-125m-e4
slm-125m-e4 is a text generation model from Ace-2504. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
A 125M-parameter LLaMA-style model for US legal/financial text, and the fourth epoch-milestone in a continued-pretraining series:
Downloads · 30 days
54
23% of all-time downloads
All-time downloads
235
Public
Parameters
126M
503 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors503 MB · 100%
From the Hugging Face model README
A 125M-parameter LLaMA-style model for US legal/financial text, and the fourth epoch-milestone in a continued-pretraining series:
slm-125m-base (v1) → extended → e2 (2 epochs on the rebuilt 2.5B corpus)
→ e4 (this model — 2 further epochs on the same corpus, 4 epochs total).
This run continued Ace-2504/slm-125m-e2 for 2 more epochs (9,440 steps, ~4.95B
tokens; ~2.47B per epoch) on the same rebuilt 2.5B-token corpus (stricter OCR
gate, exhausted case-law source). Fresh cosine schedule, peak LR 3e-4 → floor 3e-5,
one A100-40GB. It exists to answer a research question: once continued
pretraining has already recovered, do further epochs on the same in-distribution
data keep improving a small model, or start to overfit and forget?
Validation perplexity fell from e2's 9.86 to 9.44. Measured against e2's exact endpoint on the held-out forgetting probe, all three domains improved rather than regressing — the transient rise while the fresh cosine LR peaked was fully recovered and surpassed:
| Domain | loss at e2 start | loss at e4 end | change |
|---|---|---|---|
| SEC filings | 1.7470 | 1.6977 | −2.82% |
| US case law | 2.2828 | 2.2324 | −2.21% |
| FineWeb-Edu | 3.1580 | 3.0987 | −1.88% |
So on this in-distribution corpus, a second round of continued pretraining behaves
as pure improvement, not catastrophic forgetting — the opposite of what the
single-epoch out-of-distribution extended run showed on SEC text.
12 layers, 768 hidden, 12 heads, 1024 context, RoPE (θ=10000), RMSNorm, SwiGLU,
tied embeddings, vocab 16,384 (125.8M params). Same tokenizer and embedding matrix
as slm-125m-base, so token ids mean exactly what they did in v1 training.
Not instruction-tuned — this is a base model. Use Ace-2504/slm-125m-base for the
original, or Ace-2504/slm-125m-e2 for the 2-epoch checkpoint.