Skip to content

PKU-Alignment

beaver-7b-v1.0-reward

PKU-Alignment/beaver-7b-v1.0-reward

beaver-7b-v1.0-reward is a reinforcement learning model from PKU-Alignment. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for safe-rlhf.

The Beaver reward model is a preference model trained using the PKU-SafeRLHF dataset. It can play a role in the safe RLHF algorithm, helping the Beaver model become more helpful.

Downloads · 30 days

2.1K

2% of all-time downloads

All-time downloads

103K

Public

Parameters

6.6B

40.9 GB on disk

Likes

17

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors13.2 GB · 100%

Parameter types

How the weights are stored.

BF166.6B · 100%

Task
Reinforcement Learning
Library
safe-rlhf
Type
llama
Languages
en
Created
Jul 8, 2023
Updated
Apr 20, 2024