Skip to content

ligaydima

ppo-reward-model

ligaydima/ppo-reward-model

ppo-reward-model is a text classification model from ligaydima. Use it when you need a label for a piece of text. It is set up for transformers.

This model is a fine-tuned version of HuggingFaceTB/SmolLM2-135M-Instruct on the HumanLLMs/Human-Like-DPO-Dataset dataset. It has been trained using TRL.

Downloads · 30 days

10

5% of all-time downloads

All-time downloads

194

Public

Parameters

135M

10.8 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors538 MB · 99%

At a glance

Task
Text Classification
Library
transformers
Model type
llama
Access
Public
Created
Mar 6, 2025
Updated
Mar 8, 2025
SHA
9ef71a19

Try a prompt

Base models

Task
Text Classification
Library
transformers
Type
llama
Created
Mar 6, 2025
Updated
Mar 8, 2025
ppo-reward-model — AI Model — AIMarketly