Downloads ยท 30 days
59
9% of all-time downloads
Seedyai/MiniGPT-Scratch-5M
MiniGPT-Scratch-5M is a text generation model from Seedyai. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
A GPT-style language model built completely from scratch in PyTorch for educational purposes.
Downloads ยท 30 days
59
9% of all-time downloads
All-time downloads
690
Public
Parameters
5.2M
21 MB on disk
Likes
2
Trending 1
Click a slice to open those files.
.safetensors21 MB ยท 98%
From the Hugging Face model README
A GPT-style language model built completely from scratch in PyTorch for educational purposes.
This project implements the complete GPT pipeline, including tokenizer training, Transformer architecture, language model training, and text generation without relying on pre-built GPT implementations.
For installation check the repo:
https://github.com/samdoom-coder/MiniGPT.git
| Parameter | Value |
|---|---|
| Model Type | Decoder-only Transformer |
| Layers | 4 |
| Attention Heads | 4 |
| Embedding Dimension | 256 |
| Feed Forward Dimension | 1024 |
| Context Length | 128 |
| Vocabulary Size | 8000 |
| Total Parameters | ~5.2 Million |
The model is trained on:
Dataset:
https://huggingface.co/datasets/roneneldan/TinyStories
A Byte Pair Encoding (BPE) tokenizer was trained from scratch using the TinyStories dataset.
Special tokens:
<pad><bos><eos><unk>Vocabulary Size:
8000
Optimizer
AdamW
Loss Function
CrossEntropyLoss
Learning Rate
3e-4
Batch Size
16
Epochs
20
MiniGPT supports multiple decoding strategies:
Example:
output = model.generate(
input_ids,
max_new_tokens=200,
temperature=0.8,
top_k=100,
top_p=0.9,
)
Once upon a time
There was a little girl with dark hair and a smile.
She was very happy.
She ran to her mom and showed her the new dress.
Her mom smiled and said,
"Thank you, Lily.
You are a very good girl."
Lily smiled and said,
"I love you too, mom.
You are very kind."
MiniGPT
โ
โโโ src/
โ โโโ attention.py
โ โโโ config.py
โ โโโ dataset.py
โ โโโ embeddings.py
โ โโโ feedforward.py
โ โโโ model.py
โ โโโ trainer.py
โ โโโ transformer.py
โ
โโโ tokenizer/
โ โโโ tokenizer.json
โ โโโ tokenizer_config.json
โ
โโโ train.py
โโโ generate.py
โโโ train_tokenizer.py
โโโ README.md
This project was created to understand how GPT-style language models work internally by implementing every major component from scratch instead of using existing GPT implementations.
The goal is educational: to learn the architecture, training process, and text generation pipeline of modern decoder-only Transformer language models.
MIT License
โญ If you found this project interesting or helpful, consider giving it a like!