Skip to content

RLHFlow

LLaMA3-iterative-DPO-final

RLHFlow/LLaMA3-iterative-DPO-final

LLaMA3-iterative-DPO-final is a text generation model from RLHFlow. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.

Paper: RLHF Workflow: From Reward Modeling to Online RLHF (Published in TMLR, 2024) Authors: Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, Tong Zhang Code…

Downloads · 30 days

222

0% of all-time downloads

All-time downloads

69.6K

Public

Parameters

8B

16.1 GB on disk

Likes

45

Trending 2

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors16.1 GB · 100%

At a glance

Task
Text Generation
Library
transformers
License
llama3
Model type
llama
Access
Public
Created
May 17, 2024
Updated
Oct 14, 2024
SHA
8c929ad1

Try a prompt

Task
Text Generation
Library
transformers
Type
llama
License
llama3
Featherless Ai
live
Created
May 17, 2024
Updated
Oct 14, 2024