Downloads · 30 days
0
johnnycwatt/MiniStoryGPT
MiniStoryGPT is a machine learning model from johnnycwatt. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
MiniStoryGPT is a compact, educational GPT-style language model built in PyTorch to demonstrate training transformer architectures from scratch. It is trained on the TinyStories dataset to generate short, child-friend…
Downloads · 30 days
0
Access
Public
Updated Sep 22, 2025
Repo size
376 MB
Likes
0
Public
Click a slice to open those files.
.pth376 MB · 100%
From the Hugging Face model README
MiniStoryGPT is a compact, educational GPT-style language model built in PyTorch to demonstrate training transformer architectures from scratch. It is trained on the TinyStories dataset to generate short, child-friendly narratives. The model follows principles from "Attention is All You Need" and draws inspiration from Andrej Karpathy’s nanoGPT and Zero to Hero materials.
Purpose: This model is designed for educational and experimentation purposes, offering hands-on experience with building, training, and sampling from GPT-like models. It is not intended for production use.
train.bin and val.bin)MiniStoryGPT-30M.pth (367MB), saved at iteration 20,000.To use MiniStoryGPT, install the required dependencies:
pip install torch tiktoken
Download the model and mappings from this Hugging Face repository:
from huggingface_hub import hf_hub_download
hf_hub_download(repo_id="johnnycwatt/MiniStoryGPT", filename="MiniStoryGPT-30M.pth", local_dir=".")
hf_hub_download(repo_id="johnnycwatt/MiniStoryGPT", filename="old_to_new.pt", local_dir=".")
hf_hub_download(repo_id="johnnycwatt/MiniStoryGPT", filename="new_to_old.pt", local_dir=".")
Run the provided sampler.py to generate stories:
python sampler.py
Example code to load and generate:
import torch
import tiktoken
# Load model and mappings
device = 'cuda' if torch.cuda.is_available() else 'cpu'
model = GPTLanguageModel().to(device)
model.load_state_dict(torch.load("MiniStoryGPT-30M.pth", map_location=device))
old_to_new = torch.load("old_to_new.pt", map_location=device)
new_to_old = torch.load("new_to_old.pt", map_location=device)
enc = tiktoken.get_encoding("gpt2")
# Remap and generate
prompt = "Once upon a time,"
original_context = enc.encode(prompt)
remapped_context = [old_to_new.get(token, 0) for token in original_context]
context = torch.tensor([remapped_context], dtype=torch.long, device=device)
with torch.no_grad():
output = model.generate(context, max_new_tokens=300)
story = enc.decode([new_to_old.get(new_id, 0) for new_id in output[0].tolist()])
print(story)
The model requires old_to_new.pt and new_to_old.pt for token remapping due to the reduced vocabulary. See the GitHub repository for the full training and sampling code.
train.bin, val.bin) with a 10K-token vocabulary.tiktoken (GPT-2 tokenizer) with custom remapping to reduce vocab size.To reproduce training, run prepare_data.py and train.py from the GitHub repo.
Released under the MIT License. Feel free to use, modify, and distribute for research and educational purposes.
For questions or contributions, open an issue on the GitHub repository or contact [email protected]