Downloads · 30 days
0
an6308a/mini-transformer-tinystories
mini-transformer-tinystories is a text generation model from an6308a. Use it when you need the model to write or continue text. The card lists the license as mit.
A decoder-only transformer built entirely from first principles in PyTorch, no pre-built attention layers, no HuggingFace Transformers library internals. Every component (multi-head self-attention, positional encoding…
Downloads · 30 days
0
Access
Public
Updated Jul 15, 2026
Repo size
367 MB
Likes
0
Public
Click a slice to open those files.
.pt367 MB · 100%
From the Hugging Face model README
A decoder-only transformer built entirely from first principles in PyTorch, no pre-built attention layers, no HuggingFace Transformers library internals. Every component (multi-head self-attention, positional encoding, feedforward blocks, residual connections) is implemented from scratch and trained on the TinyStories dataset.
This is part of a larger project studying transformer internals, followed by implementing DeepSeek's MLA (Multi-head Latent Attention) and mHC (Manifold-Constrained Hyper-Connections) as architectural upgrades to this same baseline.
Load model.py for the architecture classes, then:
import torch
from model import MiniGPT
model = MiniGPT(vocab_size=50257, embed_dim=256, num_heads=8,
num_layers=6, hidden_dim=1024, max_seq_len=128)
ckpt = torch.load("checkpoint.pt", map_location="cpu")
model.load_state_dict(ckpt["model"])
Once upon a time, there was a little girl named Lily who loved to play outside. She had a lot of fun playing in the mud and splashing around...