Downloads · 30 days
5
9% of all-time downloads
BenBenyamin/GPT2
GPT2 is a machine learning model from BenBenyamin. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains the weights for my custom implementation of GPT-2. For the codebase, please refer to the main repository: GitHub Link
Downloads · 30 days
5
9% of all-time downloads
All-time downloads
58
Public
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.pt1.1 GB · 50%
From the Hugging Face model README
This repository contains the weights for my custom implementation of GPT-2.
For the codebase, please refer to the main repository: GitHub Link
My implementation has a different architecture structure than the original OpenAI release, but it aligns with the Hugging Face format available here:
Hugging Face GPT-2 (openai-community)
To ensure compatibility, I’ve included both the Hugging Face-compatible weights and the weights from my own training.
HF_Weights.pth
Pretrained Hugging Face GPT-2 model weights adapted to work with my implementation.
See this script for loading details:
load_gpt2_weights.py
gpt2.pth
My custom-trained GPT-2 model, trained from scratch on the FineWebEdu-10BT dataset (trained on 40BT).
To use the weights with your GPT-2 implementation, follow the example below. The model is defined here model.py as GPT2.
Below are two examples demonstrating how to load the model weights depending on whether you are using my GPT-2 or the Hugging Face-compatible weights.
import torch
from model import GPT2
model_config = {
"n_blocks": 12,
"seq_len": 1024,
"n_embd": 768,
"n_head": 12,
"vocab_size": 50304,
"dropout": 0.1 # Used for training
}
model = GPT2(**model_config)
ckpt_path = 'gpt2.pth'
state_dict = torch.load(ckpt_path, map_location='cpu')["model_state_dict"]
model.load_state_dict(state_dict)
model.eval()
import torch
from model import GPT2
hf_model_config = {
"n_blocks": 12,
"seq_len": 1024,
"n_embd": 768,
"n_head": 12,
"vocab_size": 50257,
"dropout": 0.0
}
model = GPT2(**hf_model_config)
ckpt_path = 'HF_Weights.pth'
state_dict = torch.load(ckpt_path, map_location='cpu')["model_state_dict"]
model.load_state_dict(state_dict)
model.eval()