Downloads · 30 days
12
5% of all-time downloads
glorgao/PALU-1.5B
PALU-1.5B is a text generation model from glorgao. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model is trained by PALU, a post-training strategy for concise reasoning LLMs. PALU aims to enhance reasoning efficiency by optimizing the trade-off between accuracy and output verbosity.
Downloads · 30 days
12
5% of all-time downloads
All-time downloads
230
Public
Parameters
1.8B
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
This model is trained by PALU, a post-training strategy for concise reasoning LLMs. PALU aims to enhance reasoning efficiency by optimizing the trade-off between accuracy and output verbosity.
Paper: https://arxiv.org/abs/2510.10168
Authors: Chengqian Gao, Haonan Li, Taylor W. Killian, Jianshu She, Renxi Wang, Liqun Ma, Zhoujun Cheng, Shibo Hao, and Zhiqiang Xu.
Key Features:
Concise reasoning: 65% fewer output tokens across five math benchmarks, with full reasoning preserved.
Performance: 15% higher accuracy than alternative solutions.
Efficiency: 20% faster compared with the GRPO training.
License: This repository and the model weights are released under the MIT License. The base model, DeepSeek-R1-Distill-Qwen-1.5B, is also under the MIT License, while its parent models, the Qwen-2.5 series, are licensed under the Apache 2.0 License.
Citation:
@article{gao2025concise,
title={Concise Reasoning in the Lens of Lagrangian Optimization},
author={Gao, Chengqian and Li, Haonan and Killian, Taylor W and She, Jianshu and Wang, Renxi and Ma, Liqun and Cheng, Zhoujun and Hao, Shibo and Xu, Zhiqiang},
journal={arXiv preprint arXiv:2510.10168},
year={2025}
}