Downloads · 30 days
4
16% of all-time downloads
sabersaleh/Llama2-7B-RDPO
Llama2-7B-RDPO is a machine learning model from sabersaleh. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This model is aligned using the AlpacaFarm dataset, fine-tuned through the RDPO loss. The alignment process started from the Supervised Fine-Tuned (SFT) version of LLaMA 2 7B. The optimization process was conducted wi…
Downloads · 30 days
4
16% of all-time downloads
All-time downloads
25
Public
Repo size
27 GB
Likes
0
Public
Click a slice to open those files.
.bin27 GB · 100%
From the Hugging Face model README
This model is aligned using the AlpacaFarm dataset, fine-tuned through the RDPO loss. The alignment process started from the Supervised Fine-Tuned (SFT) version of LLaMA 2 7B. The optimization process was conducted with a single epoch. For more information on the dataset, refer to the AlpacaFarm documentation (https://github.com/tatsu-lab/alpaca_farm).