Skip to content

HuangXinBa

GRPO

HuangXinBa/GRPO

GRPO is a text generation model from HuangXinBa. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.

GRPO is a causal language model fine-tuned using Group Relative Policy Optimization (GRPO), a reinforcement learning algorithm built on PPO that optimizes language models via groupwise reward comparisons. This approac…

Downloads · 30 days

24

7% of all-time downloads

All-time downloads

323

Public

Parameters

135M

538 MB on disk

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors269 MB · 98%

At a glance

Task
Text Generation
License
apache-2.0
Model type
llama
Access
Public
Created
May 27, 2025
Updated
May 28, 2025
SHA
3bbfa9a3

Try a prompt

Datasets

Task
Text Generation
Type
llama
License
apache-2.0
Languages
en
Created
May 27, 2025
Updated
May 28, 2025