Downloads · 30 days
30
24% of all-time downloads
Vibudhbh/gpt2-rlhf-implementation
gpt2-rlhf-implementation is a text generation model from Vibudhbh. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model was trained using the complete 3-stage RLHF pipeline - the same methodology used to create ChatGPT, Claude, and other state-of-the-art AI assistants!
Downloads · 30 days
30
24% of all-time downloads
All-time downloads
124
Public
Parameters
124M
498 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors498 MB · 100%
From the Hugging Face model README
This model was trained using the complete 3-stage RLHF pipeline - the same methodology used to create ChatGPT, Claude, and other state-of-the-art AI assistants!
This is a GPT-2 model that has been fine-tuned using Reinforcement Learning from Human Feedback (RLHF) with real preference data from Anthropic's HH-RLHF dataset - the same data used to train Claude.
Stage 1: Supervised Fine-Tuning (SFT)
Stage 2: Reward Model Training
Stage 3: PPO Optimization
Prompt: "How can I improve my communication skills?"
Base GPT-2: [irrelevant/confusing response]
RLHF Model: [helpful, structured advice]
Reward Score Improvement: +69.6%
from transformers import GPT2LMHeadModel, GPT2Tokenizer
# Load the model
model = GPT2LMHeadModel.from_pretrained("Vibudhbh/gpt2-rlhf-anthropic")
tokenizer = GPT2Tokenizer.from_pretrained("Vibudhbh/gpt2-rlhf-anthropic")
# Generate response
prompt = "How can I learn machine learning effectively?"
inputs = tokenizer.encode(prompt, return_tensors="pt")
with torch.no_grad():
outputs = model.generate(
inputs,
max_length=inputs.shape[1] + 50,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response[len(prompt):])
This model demonstrates:
If you use this model, please cite:
@misc{gpt2-rlhf-anthropic,
title={GPT-2 RLHF: ChatGPT-Style Training Pipeline},
author={Your Name},
year={2024},
url={https://huggingface.co/Vibudhbh/gpt2-rlhf-anthropic}
}
🚀 This model represents a complete implementation of the ChatGPT training methodology!
Built with real Anthropic data, production-grade techniques, and measurable human alignment improvements.