Skip to content

amirali1985

interpreting_reward_models

amirali1985/interpreting_reward_models

interpreting_reward_models is a machine learning model from amirali1985. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

We train a collection of models under RLHF on the above datasets. We use DPO for hh-rlhf and unalignment, and train a PPO on completing IMDB prefixes with positive sentiment.

Downloads · 30 days

0

Access

Public

Updated Aug 7, 2024

Repo size

120 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.arrow32.2 GB · 69%

At a glance

License
mit
Access
Public
Created
May 4, 2024
Updated
Aug 7, 2024
SHA
8b409eb9
License
mit
Languages
en
Created
May 4, 2024
Updated
Aug 7, 2024