Downloads · 30 days
0
amirali1985/interpreting_reward_models
interpreting_reward_models is a machine learning model from amirali1985. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
We train a collection of models under RLHF on the above datasets. We use DPO for hh-rlhf and unalignment, and train a PPO on completing IMDB prefixes with positive sentiment.
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2024
Repo size
120 GB
Likes
0
Public
Click a slice to open those files.
.arrow32.2 GB · 69%
From the Hugging Face model README
We train a collection of models under RLHF on the above datasets. We use DPO for hh-rlhf and unalignment, and train a PPO on completing IMDB prefixes with positive sentiment.