Downloads · 30 days
0
Lomesh7777/slm
slm is a text generation model from Lomesh7777. Use it when you need the model to write or continue text. The card lists the license as mit.
A ~47-51M parameter decoder-only transformer trained from scratch using a 3-stage curriculum (easy → hard text). Built entirely in PyTorch.
Downloads · 30 days
0
Access
Public
Updated Apr 7, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json6.4 MB · 98%
From the Hugging Face model README
A ~47-51M parameter decoder-only transformer trained from scratch using a 3-stage curriculum (easy → hard text). Built entirely in PyTorch.
slm/
├── model.py Transformer architecture (RoPE / learnable pos, KV cache)
├── tokenizer.py BPE tokenizer training (two vocab options)
├── dataset.py Per-stage datasets + val set builders
├── train.py Single-stage training loop (3 exit conditions + logging)
├── curriculum.py Multi-stage orchestrator
├── logger.py CSV + console logging
├── configs/
│ ├── stage0.yaml TinyStories — seq=256, max=200M tokens
│ ├── stage1.yaml SimpleWiki + BabyLM — seq=384, max=220M tokens
│ └── stage2.yaml FineWeb-Edu — seq=512, max=500M tokens
├── notebooks/
│ └── run.ipynb All training commands (run from here)
├── tokenizers/ Created after step 2
├── checkpoints/ Created during training
├── logs/ CSV logs per stage
└── cache/ Preprocessed token chunks (created automatically)
| Parameter | Value |
|---|---|
| d_model | 512 |
| n_layers | 6 |
| n_heads | 8 |
| d_ff | 2048 (SwiGLU) |
| ctx_len | 512 |
| Norm | RMSNorm |
| Activation | SwiGLU |
| Position | Learnable OR RoPE (switchable) |
| Bias | False |
| Weight tying | True (lm_head = tok_emb) |
| Params (50k tok) | ~51M (learnable) / ~49M (RoPE) |
| Stage | Dataset | Seq Len | Max Tokens | Val Source |
|---|---|---|---|---|
| 0 | TinyStories | 256 | 200M | TinyStories val |
| 1 | SimpleWiki + BabyLM | 384 | 220M | SimpleWiki sample |
| 2 | FineWeb-Edu (≥3 edu) | 512 | 500M | FineWeb sample |
All 3 val losses are logged at every eval step regardless of current stage. This lets you observe cross-stage forgetting in the logs.
patience evalspip install torch tokenizers datasets pyyaml
Or run Cell 0 in notebooks/run.ipynb.
Train both tokenizers on a sample of all stage data combined.
# Tokenizer A: corpus-derived vocab (~32-40k)
python tokenizer.py --output_dir tokenizers/ --sample_size 2_000_000 --which corpus
# Tokenizer B: fixed 50k vocab
python tokenizer.py --output_dir tokenizers/ --sample_size 2_000_000 --which fixed
# Or both at once:
python tokenizer.py --output_dir tokenizers/ --sample_size 2_000_000 --which both
Output files:
tokenizers/tokenizer_corpus.jsontokenizers/tokenizer_50k.jsonTokenizer:
tokenizers/tokenizer_50k.json — 50k vocab, larger embedding tabletokenizers/tokenizer_corpus.json — smaller natural vocab, leaner modelPositional encoding:
--pos_type learnable — standard learned position embeddings--pos_type rope — RoPE (no learned params, better length extrapolation)python train.py \
--stage 0 \
--config configs/stage0.yaml \
--tokenizer tokenizers/tokenizer_50k.json \
--pos_type learnable \
--checkpoint_dir checkpoints/ \
--log_dir logs/ \
--cache_dir cache/
Resume if interrupted:
python train.py --stage 0 --config configs/stage0.yaml \
--tokenizer tokenizers/tokenizer_50k.json \
--pos_type learnable \
--checkpoint_dir checkpoints/ --log_dir logs/ --cache_dir cache/ \
--resume
Output: checkpoints/stage0_best.pt
Loads Stage 0 weights as starting point.
python train.py \
--stage 1 \
--config configs/stage1.yaml \
--tokenizer tokenizers/tokenizer_50k.json \
--pos_type learnable \
--checkpoint_dir checkpoints/ \
--log_dir logs/ \
--cache_dir cache/ \
--prev_checkpoint checkpoints/stage0_best.pt
Output: checkpoints/stage1_best.pt
Loads Stage 1 weights as starting point.
python train.py \
--stage 2 \
--config configs/stage2.yaml \
--tokenizer tokenizers/tokenizer_50k.json \
--pos_type learnable \
--checkpoint_dir checkpoints/ \
--log_dir logs/ \
--cache_dir cache/ \
--prev_checkpoint checkpoints/stage1_best.pt
Output: checkpoints/stage2_best.pt
python curriculum.py \
--tokenizer tokenizers/tokenizer_50k.json \
--pos_type learnable \
--checkpoint_dir checkpoints/ \
--log_dir logs/ \
--cache_dir cache/
Start from a specific stage:
python curriculum.py --start_stage 1 --tokenizer tokenizers/tokenizer_50k.json ...
Run specific stages only:
python curriculum.py --stages 1 2 --tokenizer tokenizers/tokenizer_50k.json ...
# In notebook: Cell 10
# Or directly:
python -c "
import torch
from tokenizers import Tokenizer
from model import SLM
tok = Tokenizer.from_file('tokenizers/tokenizer_50k.json')
ckpt = torch.load('checkpoints/stage2_best.pt', map_location='cuda')
m = SLM(ckpt['config']).cuda()
m.load_state_dict(ckpt['model_state'])
m.eval()
ids = tok.encode('Once upon a time').ids
out = m.generate(torch.tensor([ids]).cuda(), max_new=100, temperature=0.8, top_k=50)
print(tok.decode(out[0].tolist()))
"
Run Cell 11 in notebooks/run.ipynb after training.
Each stage produces a CSV in logs/stage{N}_{timestamp}.csv:
step | stage | tokens_seen | train_loss | val_s0 | val_s1 | val_s2 | lr | note
Plot all curves with Cell 9 in the notebook:
# Or from command line:
python -c "
import glob, pandas as pd, matplotlib.pyplot as plt
# ... see notebook cell 9
"
Edit configs/stage{N}.yaml to adjust:
batch_size — reduce if OOMmax_tokens — reduce for faster experimentspatience — how many evals before plateau exitlearning_rate — per-stage LReval_interval — steps between val evaluationsTo run the baseline (random data order, same compute), train with all data mixed without staged configs. The easiest way is to:
dataset.py's _train_iter for a single stageconfigs/stage2.yaml settings (full ctx_len, total token budget = sum of all stages)The per-step CSV logs make this comparison straightforward in the notebook.
| Component | Spec |
|---|---|
| GPU | NVIDIA RTX 4060 Ti 16GB |
| VRAM used | ~2-3GB (batch=32) |
| Precision | bf16 (auto-detected) |
| Storage | ~10GB for cached data |