Skip to content

sciarrilli

Llama-3.2-3B-DPO

sciarrilli/Llama-3.2-3B-DPO

Llama-3.2-3B-DPO is a machine learning model from sciarrilli. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.

This model is a fine-tuned version of meta-llama/Llama-3.2-3B-Instruct on the trl-lib/ultrafeedbackbinarized dataset. It has been trained using TRL.

Downloads · 30 days

0

Access

Public

Updated Mar 18, 2025

Repo size

53.9 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors36.7 MB · 68%

At a glance

Library
transformers
Access
Public
Created
Mar 18, 2025
Updated
Mar 18, 2025
SHA
5cbfce95

Base models

Library
transformers
Created
Mar 18, 2025
Updated
Mar 18, 2025