Downloads · 30 days
25
1% of all-time downloads
quyanh/OpenRS-GRPO
OpenRS-GRPO is a text generation model from quyanh. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This repository hosts model for the Open RS project, accompanying the paper Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t. The project explores enhancing reasoning capabilities in sma…
Downloads · 30 days
25
1% of all-time downloads
All-time downloads
4.3K
Public
Parameters
1.8B
135 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
This repository hosts model for the Open RS project, accompanying the paper Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t. The project explores enhancing reasoning capabilities in small large language models (LLMs) using reinforcement learning (RL) under resource-constrained conditions.
We focus on a 1.5-billion-parameter model, DeepSeek-R1-Distill-Qwen-1.5B, trained on 4 NVIDIA A40 GPUs (48 GB VRAM each) within 24 hours. By adapting the Group Relative Policy Optimization (GRPO) algorithm and leveraging a curated, compact mathematical reasoning dataset, we conducted three experiments to assess performance and behavior. Key findings include:
o1-preview.These results showcase RL-based fine-tuning as a cost-effective approach for small LLMs, making reasoning capabilities accessible in resource-limited settings. We open-source our code, models, and datasets to support further research.
For more details, please refer our github.
o1-preview at 44.6%)
Our approach uses 7,000 samples (42,000 total outputs) and costs ~$42 on 4x A40 GPUs in 24 hours, compared to:
Qwen2.5-7B-SimpleRL ($1,633), Eurus-2-7B-PRIME ($1,088)DeepScaleR-1.5B-Preview ($3,629), Still-3-1.5B-Preview ($2,268)

If this project aids your work, please cite it as:
@misc{dang2025reinforcementlearningreasoningsmall,
title={Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't},
author={Quy-Anh Dang and Chris Ngo},
year={2025},
eprint={2503.16219},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.16219},
}