Downloads · 30 days
3
0% of all-time downloads
xychen123/LamPO
LamPO is a machine learning model from xychen123. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
LamPO (Lambda Policy Optimization) is a reinforcement learning framework for improving the reasoning capabilities of language models. It extends Group Relative Policy Optimization (GRPO) by replacing scalar group-mean…
Downloads · 30 days
3
0% of all-time downloads
All-time downloads
92.8K
Public
Parameters
1.2M
4.7 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors4.7 MB · 100%
From the Hugging Face model README
LamPO (Lambda Policy Optimization) is a reinforcement learning framework for improving the reasoning capabilities of language models. It extends Group Relative Policy Optimization (GRPO) by replacing scalar group-mean advantage estimation with a pairwise decomposed advantage inspired by learning-to-rank methods such as LambdaRank.
链接:论文1; 论文2
特别鸣谢:
Instead of comparing each generated response only against a group average, LambdaPO learns from fine-grained pairwise reward differences among sampled reasoning trajectories. This helps the model better distinguish high-quality reasoning paths, improve credit assignment, and reduce unstable optimization behavior during RL training.
This work is based on the paper:
“LambdaPO: A Lambda Style Policy Optimization for Reasoning Language Models” ( 链接:论文1; 论文2 )
Authors:
Corresponding author: Xinyuan Chen
@article{yuan2026lambdapo,
title={LambdaPO: A Lambda Style Policy Optimization for Reasoning Language Models},
author={Yuan, Zhe and Zhou, Yipeng and Li, Jinghan and Chen, Xinyuan and Deng, Bowen and Chen, Zhiqian and Zhao, Liang},
year={2026}
}