Downloads · 30 days
0
adimunot/transformer-from-scratch
transformer-from-scratch is a text generation model from adimunot. Use it when you need the model to write or continue text. The card lists the license as mit.
Weights for github.com/adimunot21/transformer-from-scratch, a decoder-only Transformer built from first principles in PyTorch — attention, multi-head projections, positional encoding, blocks and training loop written…
Downloads · 30 days
0
Access
Public
Updated Aug 29, 2026
Repo size
18.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors18.8 MB · 100%
From the Hugging Face model README
Weights for github.com/adimunot21/transformer-from-scratch, a decoder-only Transformer built from first principles in PyTorch — attention, multi-head projections, positional encoding, blocks and training loop written by hand, plus a BPE tokenizer also implemented from scratch.
Two models, both trained 5000 steps on Tiny Shakespeare:
| Folder | Tokenizer | Vocab | Params | d_model | Heads | Layers | Context |
|---|---|---|---|---|---|---|---|
char/ | character-level | 65 | 1.89 M | 128 | 4 | 4 | 256 |
bpe/ | BPE (from scratch) | 768 | 5.37 M | 192 | 6 | 6 | 256 |
Each folder holds model.safetensors and a config.json with the exact architecture
and training hyperparameters.
Plain state_dicts for the model class in the source repo:
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
path = hf_hub_download("adimunot/transformer-from-scratch", "bpe/model.safetensors")
model.load_state_dict(load_file(path)) # model built from config.json
model.eval()