Skip to content

LBK95

grpo-OracleReward_Async_V1

LBK95/grpo-OracleReward_Async_V1

grpo-OracleReward_Async_V1 is a machine learning model from LBK95. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.

This model is a fine-tuned version of meta-llama/Llama-3.2-1B. It has been trained using TRL.

Downloads · 30 days

0

Access

Public

Updated Jan 9, 2026

Repo size

84.9 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors22.6 MB · 57%

At a glance

Library
transformers
Access
Public
Created
Jan 8, 2026
Updated
Jan 9, 2026
SHA
82d01cce

Base models

Library
transformers
Created
Jan 8, 2026
Updated
Jan 9, 2026
grpo-OracleReward_Async_V1 — AI Model — AIMarketly