Downloads · 30 days
10
10% of all-time downloads
qingfei1/R-Search-7b-ppo
R-Search-7b-ppo is a machine learning model from qingfei1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="center" 🤗 <a href="https://huggingface.co/datasets/qingfei1/R-Searchdatasets" target="blank"[R-Search Datasets] </a • 💻 <a href="https://github.com/QingFei1/R-Search" target="blank"[Github Repo]</a </p
Downloads · 30 days
10
10% of all-time downloads
All-time downloads
105
Public
Parameters
7.6B
30.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors30.5 GB · 100%
From the Hugging Face model README
R-Search is a novel reinforcement learning framework for reasoning–search integration. It enables LLMs to autonomously perform multi-step reasoning with deep search interaction, and to learn optimal reasoning–search trajectories via multi-reward signals, substantially improving performance on complex logic- and knowledge-intensive tasks.
We open-sourced the following models trained only on the 2wikimultihopqa training set:
| Model | Huggingface Repo | Description |
|---|---|---|
| R-Search-7b-grpo | 🤗 Huggingface Repo | Trained Qwen2.5-7B-Instruct using the GRPO algorithm |
| R-Search-3b-grpo | 🤗 Huggingface Repo | Trained Qwen2.5-3B-Instruct using the GRPO algorithm |
| R-Search-7b-ppo | 🤗 Huggingface Repo | Trained Qwen2.5-7B-Instruct using the PPO algorithm |
| R-Search-3b-ppo | 🤗 Huggingface Repo | Trained Qwen2.5-3B-Instruct using the PPO algorithm |