Downloads · 30 days
0
siyzhang/EasyGPT
EasyGPT is a machine learning model from siyzhang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Dec 25, 2025
Repo size
3.6 GB
Likes
0
Public
Click a slice to open those files.
.pt3.6 GB · 100%
From the Hugging Face model README
A 303M parameter GPT-2 style model trained from scratch on the OpenWebText dataset.
Reaching a validation loss of 2.887, comparable to GPT-2 Medium.
This is a Decoder-only Transformer language model trained using Andrej Karpathy's nanoGPT framework. We integrated new components such as RMSNorm, Rotary Positional Embeddings (RoPE), SwiGLU, and GQA. It was trained from scratch on the OpenWebText dataset, which is an open-source reproduction of the dataset used to train OpenAI's GPT-2.
| Attribute | Value |
|---|---|
| Parameters | 303 Million (comparable to GPT-2 Medium) |
| Architecture | GPT-2 (1024 context window, RoPE/Standard embeddings) |
| Dataset | OpenWebText (~17GB cleaned) |
| Tokenizer | GPT-2 BPE (via tiktoken) |
| Training Steps | 15,000 steps |
| Batch Size | ~0.5M tokens per step (Gradient Accumulation) |
| Total Tokens | ~7.3 Billion tokens |
| Final Val Loss | 2.887 (PPL 18.0) |
As a Base Model (not instruction-tuned), it excels at:
Since this model is based on nanoGPT and uses a custom checkpoint format (.pt), you need the original model definition to load it.You can refer to https://github.com/ssyzhang/EasyGPT
This project is licensed under the MIT License. See the LICENSE file for the full license text.