Skip to content

RTO-RL

Llama3.2-1B-RewardModel

RTO-RL/Llama3.2-1B-RewardModel

Llama3.2-1B-RewardModel is a machine learning model from RTO-RL. Use it for the machine learning task on the model card, and read the license before you ship it in a product.

Base model: unsloth/Llama-3.2-1B-Instruct

Downloads · 30 days

44

1% of all-time downloads

All-time downloads

5.6K

Public

Parameters

1.2B

2.5 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors2.5 GB · 99%

At a glance

Model type
llama
Access
Public
Created
Nov 18, 2024
Updated
Feb 11, 2025
SHA
4fada235

Base models

Type
llama
Created
Nov 18, 2024
Updated
Feb 11, 2025