Downloads · 30 days
45
5% of all-time downloads
POLARIS-Project/Polaris-7B-Preview
Polaris-7B-Preview is a machine learning model from POLARIS-Project. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div align="center" <h1 POLARIS </h1 <div 🌠 A <strongPO</strongst-training recipe for scaling R<strongL</strong on <strongA</strongdvanced <strongR</strongeason<strongI</strongng model<strongS</strong 🚀 </div </div <br
Downloads · 30 days
45
5% of all-time downloads
All-time downloads
957
Public
Parameters
7.6B
15.2 GB on disk
Likes
32
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
Polaris is an open-source post-training method that uses reinforcement learning (RL) scaling to refine and enhance models with advanced reasoning abilities. Our research shows that even top-tier models like Qwen3-4B can achieve significant improvements on challenging reasoning tasks when optimized with Polaris. By leveraging open-source data and academic-level resources, Polaris pushes the capabilities of open-recipe reasoning models to unprecedented heights. In benchmark tests, our method even surpasses top commercial systems, including Claude-4-Opus, Grok-3-Beta, and o3-mini-high (2025/01/03).
The details of our training recipe and analysis can be found in our blog post. The code and data for reproducing our results can be found in our github repo.
| Models | AIME24 avg@32 | AIME25 avg@32 | Minerva Math avg@4 | Olympiad Bench avg@4 | AMC23 avg@8 |
|---|---|---|---|---|---|
| Deepseek-R1-Distill-Qwen-7B | 55.0 | 39.7 | 36.7 | 56.8 | 81.9 |
| AReal-boba-RL-7B | 61.9 | 48.3 | 39.5 | 61.9 | 86.4 |
| Skywork-OR1-7B-Math | 69.8 | 52.3 | 40.8 | 63.2 | 85.3 |
POLARIS-7B-Preview | 72.6 | 52.6 | 40.2 | 65.4 | 89.0 |
| Deepseek-R1-Distill-Qwen-32B | 72.6 | 54.9 | 42.1 | 59.4 | 84.3 |
| qwen3-32B | 81.4 | 72.9 | 44.2 | 66.7 | 92.4 |
| qwen3-4B | 73.8 | 65.6 | 43.6 | 62.2 | 87.2 |
POLARIS-4B-Preview | 81.2 | 79.4 | 44.0 | 69.1 | 94.8 |
The training and evaluation codebase is heavily built on Verl. The reward function in polaris in from DeepScaleR. Our model is trained on top of Qwen3-4B and DeepSeek-R1-Distill-Qwen-7B. Thanks for their wonderful work.
@misc{Polaris2025,
title = {POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models},
url = {https://hkunlp.github.io/blog/2025/Polaris},
author = {An, Chenxin and Xie, Zhihui and Li, Xiaonan and Li, Lei and Zhang, Jun and Gong, Shansan and Zhong, Ming and Xu, Jingjing and Qiu, Xipeng and Wang, Mingxuan and Kong, Lingpeng}
year = {2025}
}