Downloads · 30 days
19
0% of all-time downloads
OpenRLHF/Llama-3-8b-rm-mixture
Llama-3-8b-rm-mixture is a machine learning model from OpenRLHF. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
The Llama3-8b-based Reward Model was trained using OpenRLHF and a combination of datasets available at https://huggingface.co/datasets/OpenLLMAI/preferencedatasetmixture2andsafepku.
Downloads · 30 days
19
0% of all-time downloads
All-time downloads
52.8K
Public
Parameters
7.5B
15 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors15 GB · 100%
From the Hugging Face model README
The Llama3-8b-based Reward Model was trained using OpenRLHF and a combination of datasets available at https://huggingface.co/datasets/OpenLLMAI/preference_dataset_mixture2_and_safe_pku.
Base model: https://huggingface.co/OpenRLHF/Llama-3-8b-sft-mixture
Cosine Scheduler
Learning Rate: 9e-6
Warmup Ratio: 0.03
Batch Size: 256
Epoch: 1