Downloads · 30 days
36
1% of all-time downloads
agi-css/hh-rlhf-sft
hh-rlhf-sft is a text generation model from agi-css. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
36
1% of all-time downloads
All-time downloads
3.3K
Public
Repo size
53.9 GB
Likes
3
Public
Click a slice to open those files.
.bin27 GB · 100%
From the Hugging Face model README


Efficient, Effective, and Stable alternative of RLHF!
Instead of training an additional reward model that is likely to be gamed, we directly train the model on the social games! 🕹️ 🎲 🎮
Full details on simulation and training can be found here.
This is the second step of Stable Alignment project, which is a supervised fine-tuned model on Anthropic HH-RLHF dataset (only on the 'accepted' options).
We use the Alpaca fine-tuning script to train this model.
Although this project aims to better align current LMs with social norms, inappropriate content and inherent biases in the training data will still impair the alignment of the model.
The model should not be used directly in any application, without a prior assessment of safety and fairness concerns specific to the application.
Please cite our paper if you use the data or code in this repo:
@misc{liu2023sociallyaligned,
title={Training Socially Aligned Language Models in Simulated Human Society},
author={Ruibo Liu and Ruixin Yang and Chenyan Jia and Ge Zhang and Denny Zhou and Andrew M. Dai and Diyi Yang and Soroush Vosoughi},
year={2023},
eprint={2305.16960},
archivePrefix={arXiv},
primaryClass={cs.CL}
}