Downloads · 30 days
9
17% of all-time downloads
annaovesnaatatt/gpt2-post-ppo
gpt2-post-ppo is a feature extraction model from annaovesnaatatt. Use it when you need embeddings to search or compare text. It is set up for transformers.
This is a testing model created for RLHF. The reward model used for training is martin-arguments and was trained on 1000 examples of the Anthropic/hh-rlhf dataset.
Downloads · 30 days
9
17% of all-time downloads
All-time downloads
52
Public
Repo size
996 MB
Likes
0
Public
Click a slice to open those files.
.bin498 MB · 99%
From the Hugging Face model README
This is a testing model created for RLHF. The reward model used for training is martin-arguments and was trained on 1000 examples of the Anthropic/hh-rlhf dataset.