Downloads · 30 days
21
21% of all-time downloads
igidn/VerySmollGPT-5M-Base
VerySmollGPT-5M-Base is a text generation model from igidn. Use it when you need the model to write or continue text. The card lists the license as mit.
A lightweight character-level GPT model trained entirely on a Raspberry Pi 5. This model demonstrates that capable language models can be trained on consumer hardware with limited resources.
Downloads · 30 days
21
21% of all-time downloads
All-time downloads
99
Public
Parameters
4.8M
19.2 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors19.2 MB · 100%
From the Hugging Face model README
A lightweight character-level GPT model trained entirely on a Raspberry Pi 5. This model demonstrates that capable language models can be trained on consumer hardware with limited resources.
VerySmollGPT is a decoder-only transformer model (GPT-style architecture) designed for character-level text generation. It was trained on the TinyStories dataset to generate coherent short stories.
| Component | Value |
|---|---|
| Vocabulary Size | 104 characters |
| Embedding Dimension | 256 |
| Layers | 6 |
| Attention Heads | 8 |
| Feed-forward Dimension | 1024 |
| Context Window | 128 tokens |
| Dropout | 0.1 |
| Weight Tying | Yes (token embeddings ↔ output layer) |
Hardware:
Hyperparameters:
Training Stats:
Character-level tokenization with 104 unique tokens:
<PAD>, <UNK>, <BOS>, <EOS>pip install torch safetensors
from safetensors.torch import load_file
import torch
import torch.nn as nn
# Load model weights
state_dict = load_file('model.safetensors')
# Load configuration
import json
with open('config.json', 'r') as f:
config = json.load(f)
# Note: You'll need to implement the VerySmollGPT architecture
# or use the original model.py from the repository
# Assuming you have the model loaded
model.eval()
# Encode your prompt (character-level)
prompt = "Once upon a time"
input_ids = [char_to_idx[c] for c in prompt]
input_tensor = torch.tensor([input_ids], dtype=torch.long)
# Generate
with torch.no_grad():
output_ids = model.generate(
input_tensor,
max_new_tokens=200,
temperature=0.8,
top_k=40
)
# Decode output
generated_text = ''.join([idx_to_char[i] for i in output_ids[0].tolist()])
print(generated_text)
Prompt: "Once upon a time"
Generated:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite was a penguin that had a shiny metal box on it. Timmy liked to...
Prompt: "The quick brown fox"
Generated:
The quick brown fox wanted to play with him again. The fox said he was not fair anymore. He said he was sorry and that he learned his lesson...
This model was intentionally trained on a Raspberry Pi 5 to demonstrate low-power AI training: