Downloads ยท 30 days
0
Shiv-22/tinylm
tinylm is a text generation model from Shiv-22. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as apache-2.0.
A 275M parameter small language model trained from scratch with Multi-head Latent Attention (MLA, DeepSeek-V2 style) and the Muon optimizer, on 8B unique tokens of FineWeb-Edu. Benchmarked against TinyLlama-1.1B as paโฆ
Downloads ยท 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.pt2.3 GB ยท 100%
From the Hugging Face model README
A 275M parameter small language model trained from scratch with Multi-head Latent Attention (MLA, DeepSeek-V2 style) and the Muon optimizer, on 8B unique tokens of FineWeb-Edu. Benchmarked against TinyLlama-1.1B as part of a 4-arm architecture ablation.
results/FINDINGS.mdresults/hpc_rerun_ablation.md in the repoShiv-22/tinylm-checkpoints-v2This repo holds the Run D arm of the ablation (MLA + Muon) โ the best-performing of the four and the model intended for downstream use.
| Repo | What it is |
|---|---|
Shiv-22/tinylm โ this repo | Base 275M โ Run D (MLA + Muon), ablation winner; the model for downstream use |
Shiv-22/tinylm-instruct | Instruct โ this base SmolTalk-SFT'd for chat (ChatML) |
Shiv-22/tinylm-checkpoints-v2 | All 4 ablation arms (A/B/C/D) + the E3-full continued-pretraining checkpoints |
Shiv-22/tinylm-checkpoints | v1 historical checkpoint (1Bร21 tokens, pre data-fix) |
Source & full results: github.com/shivnarainms22/TinyLM
| Parameters | 274.6M |
| Architecture | TinyLlama-style decoder-only Transformer with MLA |
| Layers | 18 |
| Hidden size | 1024 |
| Attention heads | 16 |
| MLA latent dim | 512 (decoupled RoPE 64) |
| FFN hidden | 2816 (SwiGLU) |
| Context length | 2048 |
| Vocab | 32,000 |
| Tokenizer | meta-llama/Llama-2-7b-hf |
| Tied embeddings | Yes |
| Precision | bfloat16 |
KV-cache reduction (MLA). At inference MLA caches a compressed latent plus a
decoupled RoPE key โ d_latent + d_rope = 576 values per token per layer โ instead
of an equivalent MHA's full K and V (2ยทd_model = 2048). That is a 3.56ร smaller
KV cache (71.9% reduction): 144.0 MiB โ 40.5 MiB at full 2048-token context.
Verified by test_kv_cache_shape_incremental; derivation in
results/kv_cache_reduction.md.
| Dataset | HuggingFaceFW/fineweb-edu (8B unique tokens) |
| Tokens processed | ~24B (~3 epochs) |
| Steps | 23,000 (warmup 2,000) |
| Effective batch | 512 sequences ร 2048 tokens โ 1.05M tokens/step |
| Optimizer | Muon for matrix params (lr 0.02) + AdamW for scalar/embed/LM-head/LN (lr 0.001, wd 0.1) |
| LR schedule | Cosine with linear warmup |
| Grad clip | 1.0 |
| Hardware | A100-40GB (Northeastern Explorer HPC) |
| Codebase base | modded-nanogpt |
Pure FineWeb-Edu throughout (no annealing mix, no instruction tuning).
Training logs: loss curves and throughput for this run (and all four ablation arms) are
public on Weights & Biases โ tinylm
(this model is run_D_mla_muon_v2).
0-shot eval via lm-evaluation-harness. HellaSwag and ARC-Easy reported as acc_norm (length-normalized accuracy); LAMBADA and Winogrande as acc.
| Benchmark | Metric | TinyLM 275M (this model) | TinyLlama-1.1B baseline | ฮ |
|---|---|---|---|---|
| HellaSwag | acc_norm | 41.23% | 59.1% | โ17.9 |
| ARC-Easy | acc_norm | 51.22% | 55.7% | โ4.5 |
| LAMBADA | acc | 36.81% | 58.9% | โ22.1 |
| Winogrande | acc | 51.30% | 58.9% | โ7.6 |
| Average | 45.14% | 58.2% | โ13.1 |
Baseline = TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T.
The four arms differ only in attention class and matrix optimizer; all other training settings are identical.
| AdamW | Muon | ฮ (Muon โ AdamW) | |
|---|---|---|---|
| MHA | Run A: 43.62% | Run C: 44.63% | +1.02 |
| MLA | Run B: 44.11% | Run D: 45.14% (this model) | +1.03 |
| ฮ (MLA โ MHA) | +0.49 | +0.50 | โ |
Findings:
HellaSwag and LAMBADA are the cleanest signals (monotonic A < B < C < D); ARC-Easy and Winogrande sit within stderr across all four arms.
This model uses a custom architecture (MLA with decoupled RoPE) that is not
in HuggingFace transformers. To load it, install the source repo:
git clone https://github.com/shivnarainms22/TinyLM
cd TinyLM
pip install torch transformers huggingface_hub
Then:
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
from tinylm.model import TinyLM, ModelConfig
# Download checkpoint
ckpt_path = hf_hub_download(repo_id="Shiv-22/tinylm", filename="step_22999.pt")
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=True)
# Build and load model
model = TinyLM(ModelConfig(**ckpt["config"]))
state = ckpt["model"]
if any(k.startswith("_orig_mod.") for k in state):
state = {k.removeprefix("_orig_mod."): v for k, v in state.items()}
model.load_state_dict(state)
model.eval().to("cuda").to(torch.bfloat16)
# Tokenize and generate
tok = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-hf")
prompt_ids = tok.encode("The capital of France is", return_tensors="pt").to("cuda")
with torch.no_grad():
for _ in range(20):
logits = model(prompt_ids)
next_id = logits[0, -1].argmax(dim=-1, keepdim=True).unsqueeze(0)
prompt_ids = torch.cat([prompt_ids, next_id], dim=1)
print(tok.decode(prompt_ids[0]))
acc_norm metrics; differences smaller than that should be read as noise.transformers.AutoModel.from_pretrained out of the box.v1 (2026-05, RunPod A100-80GB): single MLA+Muon training run on 1B unique tokens repeated ~21ร. Established the training pipeline but the data looping hurt long-range coherence (LAMBADA acc 29.2%). v1 weights preserved at Shiv-22/tinylm-checkpoints for contrast.
HPC re-run (2026-05, Northeastern Explorer A100-40GB): full 4-arm ablation on 8B unique tokens. The weights in this repo are from the Run D arm of that re-run.
v2 (2026-06/07): continued-pretraining probes off this model testing whether better data improves it โ fresh tokens, a web/code/math mixture, and a distilled-text recipe scaled to 7.3B tokens. Language modeling improved substantially (LAMBADA ppl 26.54 โ 23.20); commonsense reasoning did not move.
v3 (2026-07): diagnostics on that negative result (few-shot, held-out code/math perplexity โ 4.2ร / 2.0ร lower, invisible to the benchmark suite) plus a SmolTalk SFT, published as Shiv-22/tinylm-instruct.
v4 (2026-07): logit knowledge distillation from TinyLlama-1.1B into this model, as a controlled contrast against a cross-entropy control. No reasoning transfer, and language modeling got worse (ppl 26.54 โ 28.85).
Re-running the same MLA+Muon arm with the data fix (1Bร21 โ 8B unique) was worth +3.97 pts average โ roughly 2.6ร the architecture-and-optimizer ablation gain. Data quality dominates architecture at this scale.
The project's headline finding: across three independent interventions โ more data,
better data, and distillation from a 4ร larger teacher โ commonsense reasoning never moved,
while every one of them improved language modeling. At 275M, reasoning appears to be
capacity-bound. Full evidence and limits in
results/FINDINGS.md.
@misc{tinylm-275m,
author = {Shivnarain},
title = {TinyLM 275M: A small language model with MLA and Muon},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/Shiv-22/tinylm},
}
Apache 2.0. Inherits the permissive terms of modded-nanogpt (MIT) for the codebase and FineWeb-Edu (ODC-By) for the training data.