Downloads · 30 days
22
0% of all-time downloads
JoeYing/ReTool-Qwen-32B
ReTool-Qwen-32B is a machine learning model from JoeYing. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
In this work, we embrace the RL paradigm and introduce ReTool, a Tool-augmented Reinforcement learning framework explicitly designed to guide LLMs towards optimal strategies for leveraging external computational tools…
Downloads · 30 days
22
0% of all-time downloads
All-time downloads
7.4K
Public
Parameters
32.8B
131 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors131 GB · 100%
From the Hugging Face model README
In this work, we embrace the RL paradigm and introduce ReTool, a Tool-augmented Reinforcement learning framework explicitly designed to guide LLMs towards optimal strategies for leveraging external computational tools during reasoning. Our comprehensive experiments on AIME2024 and AIME2025 demonstrate that ReTool not only achieves superior accuracy compared to conventional text-based RL approaches, but also converges with significantly fewer training steps.
🚀 ReTool achieves accuracy of 67.0% on AIME 2024 and 49.3% on AIME 2025 based on the Qwen2.5-32B-Instruct model, outperforming the text-based RL baseline with less than 50% training steps.
If you find our project helpful, please cite:
@misc{feng2025retoolreinforcementlearningstrategic,
title={ReTool: Reinforcement Learning for Strategic Tool Use in LLMs},
author={Jiazhan Feng and Shijue Huang and Xingwei Qu and Ge Zhang and Yujia Qin and Baoquan Zhong and Chengquan Jiang and Jinxin Chi and Wanjun Zhong},
year={2025},
eprint={2504.11536},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2504.11536},
}