Downloads · 30 days
0
noOneCodes/tinystories-gpt-124m
tinystories-gpt-124m is a machine learning model from noOneCodes. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A 124M parameter GPT-2 architecture model implemented from scratch (following Sebastian Raschka's Build a Large Language Model from Scratch) and pretrained on the TinyStories dataset.
Downloads · 30 days
0
Access
Public
Updated Jul 19, 2026
Repo size
653 MB
Likes
0
Public
Click a slice to open those files.
.pth653 MB · 100%
From the Hugging Face model README
A 124M parameter GPT-2 architecture model implemented from scratch (following Sebastian Raschka's Build a Large Language Model from Scratch) and pretrained on the TinyStories dataset.
This is not using the prebuilt transformers library— it uses a custom GPTModel class,
included as CodingGPTFr.py.
```python import torch from CodingGPTFr import GPTModel
config = {"vocab_size": 50257, "context_length": 256, "emb_dim": 768, "n_heads": 12, "n_layers": 12, "drop_rate": 0.1, "qkv_bias": False} model = GPTModel(config) model.load_state_dict(torch.load("model_final.pth", map_location="cpu")) model.eval() ```
See generate.py for sampling with temperature and top-k.
Trained only on synthetic children's stories. It writes fluent, simple narratives but has no factual knowledge, cannot do arithmetic, and is a base model (not instruction-tuned) — it completes text rather than answering questions. Logical consistency across a story degrades at this scale.