Downloads ยท 30 days
120
4% of all-time downloads
Gen-Verse/DemyAgent-4B
DemyAgent-4B is a machine learning model from Gen-Verse. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads ยท 30 days
120
4% of all-time downloads
All-time downloads
2.9K
Public
Parameters
4.4B
8.8 GB on disk
Likes
11
Public
Click a slice to open those files.
.safetensors8.8 GB ยท 100%
From the Hugging Face model README
This repository contains the DemyAgent-4B model weights, a 4B-sized agentic reasoning model that achieves state-of-the-art performance on challenging benchmarks including AIME2024/2025, GPQA-Diamond, and LiveCodeBench-v6. DemyAgent-4B is trained using our GRPO-TCR recipe with 30K high-quality agentic RL data, demonstrating that small models can outperform much larger alternatives (14B/32B) through effective RL training strategies.
In our work, we systematically investigate three dimensions of agentic RL: data, algorithms, and reasoning modes. Our findings reveal:
| Type | Name | Link |
|---|---|---|
| ๐ Dataset | 3K Agentic SFT Data | ๐ค HuggingFace |
| ๐ Dataset | 30K Agentic RL Data | ๐ค HuggingFace |
| ๐ค Model | Qwen2.5-7B-RA-SFT | ๐ค HuggingFace |
| ๐ค Model | Qwen3-4B-RA-SFT | ๐ค HuggingFace |
| ๐ค Model | DemyAgent-4B | ๐ค HuggingFace |
Note:
- Qwen2.5-7B-RA-SFT and Qwen3-4B-RA-SFT are finetuned from Qwen2.5-7B-Instruct and Qwen3-4B-Instruct-2507 using our 3K Agentic SFT Data
- DemyAgent-4B is trained through Agentic RL with our 30K Agentic RL data using the GRPO-TCR recipe
We evaluate our models on challenging benchmarks spanning mathematics, science, and code generation tasks.
| MATH | Science | Code | ||
|---|---|---|---|---|
| Method | AIME2024 | AIME2025 | GPQA-Diamond | LiveCodeBench-v6 |
| Self-Contained Reasoning | ||||
| Qwen2.5-7B-Instruct | 16.7 | 10.0 | 31.3 | 15.2 |
| Qwen3-4B-Instruct-2507 | 63.3 | 47.4 | 52.0 | 35.1 |
| Qwen2.5-72B-Instruct | 18.9 | 15.0 | 49.0 | - |
| DeepSeek-V3 | 39.2 | 28.8 | 59.1 | 16.1 |
| DeepSeek-R1-Distill-32B | 70.0 | 46.7 | 59.6 | - |
| DeepSeek-R1-Zero (671B) | 71.0 | 53.5 | 59.6 | - |
| Agentic Reasoning | ||||
| Qwen2.5-7B-Instruct | 4.8 | 5.6 | 25.5 | 12.2 |
| Qwen3-4B-Instruct-2507 | 17.9 | 16.3 | 44.3 | 23.0 |
| ToRL-7B | 43.3 | 30.0 | - | - |
| ReTool-32B | 72.5 | 54.3 | - | - |
| Tool-Star-3B | 20.0 | 16.7 | - | - |
| ARPO-7B | 30.0 | 30.0 | 53.0 | 18.3 |
| rStar2-Agent-14B | 80.6 | <u>69.8</u> | 60.9 | - |
| DemyAgent-4B (Ours) | <u>72.6</u> | 70.0 | <u>58.5</u> | <u>26.8</u> |
โจ Despite having only 4B parameters, DemyAgent-4B achieves:
@article{yu2025demystify,
title={Demystifying Reinforcement Learning in Agentic Reasoning},
author={Yu, Zhaochen and Yang, Ling and Zou, Jiaru and Yan, Shuicheng and Wang, Mengdi},
journal={arXiv preprint arXiv:2510.11701},
year={2025}
}