Skip to content

PKU-Alignment

beaver-7b-v2.0-cost

PKU-Alignment/beaver-7b-v2.0-cost

beaver-7b-v2.0-cost is a reinforcement learning model from PKU-Alignment. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for safe-rlhf.

The Beaver cost model is a preference model trained using the PKU-SafeRLHF dataset. It can play a role in the safe RLHF algorithm, helping the Beaver model become more safe and harmless.

Downloads · 30 days

20

2% of all-time downloads

All-time downloads

1.3K

Public

Parameters

6.6B

14.5 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors13.2 GB · 100%

Parameter types

How the weights are stored.

BF166.6B · 100%

Task
Reinforcement Learning
Library
safe-rlhf
Type
llama
Languages
en
Created
Apr 19, 2024
Updated
Apr 20, 2024