Downloads · 30 days
7
41% of all-time downloads
flbjcy/LM-from-scratch
LM-from-scratch is a machine learning model from flbjcy. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A small Transformer language model, implemented entirely from scratch (no torch.nn layers beyond Parameter and containers) as part of Stanford's CS336 (Language Modeling from Scratch) assignment.
Downloads · 30 days
7
41% of all-time downloads
All-time downloads
17
Public
Repo size
91 MB
Likes
0
Public
Click a slice to open those files.
.pt90.8 MB · 100%
From the Hugging Face model README
A small Transformer language model, implemented entirely from scratch (no torch.nn layers beyond Parameter and containers) as part of Stanford's CS336 (Language Modeling from Scratch) assignment.
This model was trained only on TinyStories — a dataset of simple children's stories. It is a base language model, not an instruction-tuned assistant. It can generate coherent, story-like text continuations (especially when prompted with story-like openings such as "Once upon a time..."), but it has no capability for question-answering, factual knowledge, instruction-following, or conversation. Prompting it with anything outside the TinyStories domain will typically produce fluent but nonsensical output, since it has only ever learned to continue text in that narrow style.
This model uses a custom architecture (not directly compatible with transformers' AutoModel). To use it, you'll need the original implementation from this repo.
import torch, pickle, json
from cs336_basics.model import TransformerLM, decode
from cs336_basics.tokenizer import Tokenizer
with open("config.json") as f:
config = json.load(f)
with open("vocab.pkl", "rb") as f:
vocab = pickle.load(f)
with open("merges.pkl", "rb") as f:
merges = pickle.load(f)
tokenizer = Tokenizer(vocab, merges, special_tokens=["<|endoftext|>"])
model = TransformerLM(**config)
model.load_state_dict(torch.load("model_weights.pt"))
output = decode(model, tokenizer, "Once upon a time", max_tokens=100, temperature=0.8, top_p=0.9)
print(output)
Trained from scratch on a MacBook Air (Apple Silicon, MPS), using a custom implementation of AdamW, cosine learning rate scheduling with warmup, and gradient clipping. See the source repository for full implementation details.