Downloads · 30 days
20
28% of all-time downloads
mickey5k/slm-125m-base
slm-125m-base is a text generation model from mickey5k. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A 125M-parameter, Llama-style, decoder-only language model, pretrained completely from scratch (fresh weights, fresh 16K byte-level BPE tokenizer) on a domain-specific corpus of US case law, SEC filings, and general w…
Downloads · 30 days
20
28% of all-time downloads
All-time downloads
71
Public
Parameters
126M
252 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors252 MB · 100%
From the Hugging Face model README
A 125M-parameter, Llama-style, decoder-only language model, pretrained completely from scratch (fresh weights, fresh 16K byte-level BPE tokenizer) on a domain-specific corpus of US case law, SEC filings, and general web text.
This is a base completer, not a chat model — give it a passage prefix and it continues it. It has not been instruction-tuned, aligned, or fine-tuned for any downstream task.
| Parameters | ~125M |
| Layers | 12 |
| Hidden size | 768 |
| Attention heads | 12 (head dim 64), plain MHA |
| MLP | SwiGLU, inner 3072 |
| Normalization | RMSNorm, pre-norm |
| Position embeddings | RoPE |
| Context length | 1024 |
| Vocab | 16,384 (fresh byte-level BPE, no OOV token ever) |
| Embeddings | tied |
~3.53B packed tokens across three sources, deduplicated (MinHash-LSH on the legal corpus) and decontaminated against the LexGLUE/CaseHOLD held-out splits:
| Source | Tokens | Share |
|---|---|---|
HFforLegal/case-law (US case law, us split) | ~1.1B | 31% |
PleIAs/SEC (SEC filings) | ~2.0B | 57% |
HuggingFaceFW/fineweb-edu (general web text) | ~1.0B | 28% |
(Shares don't sum to 100% due to rounding; case-law's split was capped at its
full-dataset ceiling, not a target percentage — see the build's docs/01-data.md.)
AdamW (β 0.9/0.95, weight decay 0.1), linear warmup + cosine LR, bf16 (autocast, fp32 master weights), single-node DDP across 8x H100. Trained to step 7,500 (~4.1B tokens processed, more than one full epoch over the training set) before the run was stopped.
Held-out validation perplexity: 10.14 (full validation set, 34,513 windows). This is the honest headline metric at this scale — not MMLU or similar benchmarks, which are near-random for a 125M model.
Perplexity varies a lot by domain, reflecting the training mix:
| Domain | Perplexity |
|---|---|
| SEC filings | ~6.0 |
| Case law | ~15.3 |
| General web text (fineweb-edu) | ~27.4 |
The model is noticeably better at the two legal/financial domains it mostly trained on than at general web text.
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("mickey5k/slm-125m-base")
model = AutoModelForCausalLM.from_pretrained("mickey5k/slm-125m-base")
prompt = "The court held that"
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
output_ids = model.generate(
input_ids,
max_new_tokens=100,
min_new_tokens=20, # a base model can emit EOS almost immediately otherwise
do_sample=True,
temperature=0.8,
top_p=0.9,
top_k=50,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
docs/05-pretrain.md for the full training log and the reasoning behind
stopping at step 7,500.