Downloads · 30 days
17
29% of all-time downloads
maximehip/small-stories
small-stories is a text generation model from maximehip. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A lightweight GPT model (23M parameters) pre-trained on the TinyStories dataset for text generation.
Downloads · 30 days
17
29% of all-time downloads
All-time downloads
58
Public
Repo size
120 MB
Likes
0
Public
Click a slice to open those files.
.pt120 MB · 100%
From the Hugging Face model README
A lightweight GPT model (23M parameters) pre-trained on the TinyStories dataset for text generation.
This is a small-scale GPT (Generative Pre-trained Transformer) model trained from scratch on the TinyStories dataset. The model is designed to be efficient and suitable for deployment on resource-constrained devices while maintaining good text generation capabilities.
Model Type: Causal Language Model (Auto-regressive) Architecture: GPT-2 style decoder-only transformer Parameters: ~23 million Context Length: 256 tokens Vocabulary Size: 50,257 (GPT-2 tokenizer)
- Layers: 6 transformer blocks
- Attention Heads: 6 heads per layer
- Embedding Dimension: 384
- Hidden Dimension (MLP): 1,536 (4x expansion)
- Activation: GELU
- Normalization: LayerNorm
- Position Embeddings: Learned absolute positions
- Weight Tying: Shared embeddings between input and output
| Parameter | Value |
|---|---|
| Training Steps | 100,000 |
| Batch Size | 128 |
| Gradient Accumulation | 2 steps (effective batch size: 256) |
| Learning Rate | 6e-4 |
| LR Schedule | Linear warmup (2,000 steps) + Cosine decay |
| Min LR | 6e-5 |
| Optimizer | AdamW (β1=0.9, β2=0.95, ε=1e-9) |
| Weight Decay | 0.1 |
| Gradient Clipping | 0.5 |
| Dropout | 0.1 |
| Precision | bfloat16 (mixed precision) |
| Context Length | 256 tokens |
pip install torch tiktoken
import torch
import tiktoken
from model import GPT, GPTConfig
# Load tokenizer
enc = tiktoken.get_encoding("gpt2")
# Model configuration
config = GPTConfig(
vocab_size=50257,
block_size=256,
n_layer=6,
n_head=6,
n_embd=384,
dropout=0.0,
bias=True
)
# Load model
device = "cuda" if torch.cuda.is_available() else "cpu"
model = GPT(config).to(device)
model.eval()
# Load checkpoint
checkpoint = torch.load("pretrained_tinystories.pt", map_location=device)
checkpoint = {k.replace("_orig_mod.", ""): v for k, v in checkpoint.items()}
model.load_state_dict(checkpoint)
# Generate text
prompt = "Once upon a time, there was a little girl"
context = torch.tensor(enc.encode(prompt)).unsqueeze(0).to(device)
with torch.no_grad():
output = model.generate(
context,
max_new_tokens=200,
temperature=0.8,
top_k=40
)
generated_text = enc.decode(output[0].tolist())
print(generated_text)
✅ Generates coherent short stories in the style of children's tales ✅ Maintains context over 256 tokens ✅ Good grammar and sentence structure ✅ Fast inference (~50-100 tokens/second on GPU) ✅ Small model size (23MB) suitable for edge deployment
⚠️ Limited to simple, child-like narratives (trained on TinyStories) ⚠️ May generate repetitive text with high temperatures ⚠️ Context window limited to 256 tokens ⚠️ Not suitable for complex reasoning or factual accuracy ⚠️ English only
This model can be fine-tuned on custom datasets. Example use case: fine-tuning on Harry Potter books for story generation in that style.
See finetune.py for the fine-tuning script.
If you use this model, please cite:
@misc{tinystories-gpt-slm,
author = {maximehip},
title = {SmallStories},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/maximehip/small-stories}},
}
And the original TinyStories dataset:
@article{eldan2023tinystories,
title={TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
author={Eldan, Ronen and Li, Yuanzhi},
journal={arXiv preprint arXiv:2305.07759},
year={2023}
}
This model is released under the MIT License.
The TinyStories dataset is released under the CDL-Permissive-2.0 license.