Downloads · 30 days
17
18% of all-time downloads
dyyyyyyyy/FAPO-32B
FAPO-32B is a machine learning model from dyyyyyyyy. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This Model is trained on the FAPO-Reasoning-Dataset with generative rewards by FAPO-GenRM-4B.
Downloads · 30 days
17
18% of all-time downloads
All-time downloads
93
Public
Parameters
32.8B
65.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors65.5 GB · 100%
From the Hugging Face model README
This Model is trained on the FAPO-Reasoning-Dataset with generative rewards by FAPO-GenRM-4B.
Project Homepage: https://fapo-rl.github.io/
Code Implementation: https://github.com/volcengine/verl/tree/main/recipe/fapo
Welcome to follow and cite our works!
BibTeX citation:
@article{ding2025fapo,
title={FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning},
author={Ding, Yuyang and Zhang, Chi and Li, Juntao and Lin, Haibin and Liu, Xin and Zhang, Min},
journal={arXiv preprint arXiv:2510.22543},
year={2025}
}