Skip to content

PKU-Alignment

beaver-7b-v2.0

PKU-Alignment/beaver-7b-v2.0

beaver-7b-v2.0 is a reinforcement learning model from PKU-Alignment. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for safe-rlhf.

Beaver is a chat assistant trained based on the Standord Alpaca model (reproduced version) using the PKU-Alignment/safe-rlhf library.

Downloads · 30 days

20

3% of all-time downloads

All-time downloads

731

Public

Parameters

6.7B

13.5 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors13.5 GB · 100%

At a glance

Task
Reinforcement Learning
Library
safe-rlhf
Model type
llama
Access
Public
Created
Apr 19, 2024
Updated
May 9, 2024
SHA
05adcfbc
Task
Reinforcement Learning
Library
safe-rlhf
Type
llama
Languages
en
Created
Apr 19, 2024
Updated
May 9, 2024
beaver-7b-v2.0 — AI Model — AIMarketly