Downloads · 30 days
24
7% of all-time downloads
HuangXinBa/GRPO
GRPO is a text generation model from HuangXinBa. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
GRPO is a causal language model fine-tuned using Group Relative Policy Optimization (GRPO), a reinforcement learning algorithm built on PPO that optimizes language models via groupwise reward comparisons. This approac…
Downloads · 30 days
24
7% of all-time downloads
All-time downloads
323
Public
Parameters
135M
538 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors269 MB · 98%
From the Hugging Face model README
GRPO is a causal language model fine-tuned using Group Relative Policy Optimization (GRPO), a reinforcement learning algorithm built on PPO that optimizes language models via groupwise reward comparisons. This approach aligns model outputs with reward functions through relative ranking among multiple completions per prompt, making it well-suited for structured generation tasks such as Chain-of-Thought (CoT) reasoning.
HuggingFaceTB/SmolLM-135M-Instructreward_len (length-based)from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("HuangXinBa/GRPO")
tokenizer = AutoTokenizer.from_pretrained("HuangXinBa/GRPO")
prompt = "<your prompt here>"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
@misc{grpo2025,
title={GRPO-1: Finetuning a Language Model with Generalized Reinforcement Policy Optimization},
author={Huang Jinting},
year={2025},
note={https://wandb.ai/ggg7334-the-school-of-the-new-york-times/GRPO}
}