Skip to content

LlameUser

Qwen2.5-3B-Open-R1-GRPO

LlameUser/Qwen2.5-3B-Open-R1-GRPO

Qwen2.5-3B-Open-R1-GRPO is a text generation model from LlameUser. Use it when you need the model to write or continue text. It is set up for transformers.

This model is a fine-tuned version of Qwen/Qwen2.5-3B-Instruct on the open-r1/Big-Math-RL-Verified-Processed dataset. It has been trained using TRL.

Downloads · 30 days

21

17% of all-time downloads

All-time downloads

122

Public

Parameters

3.1B

6.2 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors6.2 GB · 100%

At a glance

Task
Text Generation
Library
transformers
Model type
qwen2
Access
Public
Created
Aug 27, 2025
Updated
Sep 1, 2025
SHA
4cb95380

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
qwen2
Created
Aug 27, 2025
Updated
Sep 1, 2025