Downloads · 30 days
0
karthik-2905/model-a-scratch
model-a-scratch is a text generation model from karthik-2905. Use it when you need the model to write or continue text. It is set up for numpy. The card lists the license as cc-by-nc-4.0.
A 3.87M-parameter decoder-only language model built entirely from scratch in pure NumPy — no PyTorch, no JAX, no autograd library. Custom reverse-mode autograd, tokenizer, training loop, KV-cache, and sampler are all…
Downloads · 30 days
0
Access
Public
Updated Jul 15, 2026
Repo size
102 MB
Likes
0
Public
Click a slice to open those files.
.npz92.9 MB · 71%
From the Hugging Face model README
A 3.87M-parameter decoder-only language model built entirely from scratch in pure NumPy — no PyTorch, no JAX, no autograd library. Custom reverse-mode autograd, tokenizer, training loop, KV-cache, and sampler are all hand-written. Optional CuPy backend swap (ZYN_BACKEND=cuda) for GPU training.
Scope is deliberately narrow: short English small-talk / companion replies only. No code generation, no tools, no retrieval, no function-calling — by design.
| Component | Choice |
|---|---|
| Positions | RoPE (rotary) |
| Norm | RMSNorm, Pre-LN |
| Attention | Grouped-Query Attention (8 query heads, 2 KV heads) + QK-Norm |
| MLP | SwiGLU (2/3 hidden-dim rule) |
| Head | Weight-tied to token embedding |
| Tokenizer | Byte-level BPE, vocab 4096, chat special tokens |
| Layers / d_model / head_dim | 4 / 256 / 32 |
| Context length | 256 |
| Params | 3,869,184 |
| Dtype | float64 (CPU / gradcheck), float32 (GPU) |
Pretrain — DailyDialog (human everyday dialogue), formatted <bos><|user|>…<eos><|assistant|>…<eos>.
Chat fine-tune (SFT) — EmpatheticDialogues (warm, supportive human dialogue) with loss masking (only assistant turns are supervised; user tokens use ignore_index).
| Path | What |
|---|---|
mla/ | Core library: tensor.py (autograd), model.py, tokenizer.py, kvcache.py, generate.py, chat.py, optim.py, loss.py |
scripts/ | pretrain.py, finetune_sft.py, build_sft_corpus.py, tokenize_sft.py, evaluate.py |
checkpoints/pretrain_final.npz | Base pretrained model |
checkpoints/sft_final.npz | Fine-tuned companion model (use this) |
data/tokenizer/tokenizer.json | Byte-BPE tokenizer |
tests/ | Gradcheck, KV-cache equivalence, sampling, chat-runtime tests |
from mla.checkpoint import load_checkpoint
from mla.tokenizer import Tokenizer
from mla.chat import ChatSession
tok = Tokenizer.load("data/tokenizer/tokenizer.json")
model, _, _ = load_checkpoint("checkpoints/sft_final.npz")
chat = ChatSession(model, tok, temperature=0.8, top_k=40, top_p=0.9)
print(chat.reply("I had a rough day today."))
Inference features: greedy / temperature / top-k / top-p sampling, KV-cache (numerically identical to full forward, verified in tests), multi-turn chat runtime, optional system= persona conditioning.
Built atom-by-atom (Karpathy style) to understand every layer of a modern LLM from first principles — each op gradchecked, each stage gated (overfit, KV-cache logit-equivalence, checkpoint resume) before moving on.