Downloads · 30 days
11
22% of all-time downloads
yanyc/InftyThink-Plus-TE-4B
InftyThink-Plus-TE-4B is a machine learning model from yanyc. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<div align="center" <h1InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning</h2 </div
Downloads · 30 days
11
22% of all-time downloads
All-time downloads
49
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
Yuchen Yan<sup>1,2,*</sup>, Liang Jiang<sup>2</sup>, Jin Jiang<sup>3</sup>, Shuaicheng Li<sup>2</sup>, <br> Zujie Wen<sup>2</sup>, Zhiqiang Zhang<sup>2</sup>, Jun Zhou<sup>2</sup>, Jian Shao<sup>1,†</sup>, Yueting Zhaung<sup>1</sup>, Yongliang Shen<sup>1,†</sup>
<sup>1</sup>Zhejiang University,
<sup>2</sup>Ant Group,
<sup>3</sup>Peking University
<em>ICML 2026</em>
<sup>*</sup>Contribution during internship at Ant Group. <sup>†</sup>Corresponding Author
Building upon our previous work <a href='https://github.com/ZJU-REAL/InftyThink'>InftyThink</a>, we introduce InftyThink+, an end-to-end reinforcement learning framework that directly optimizes the complete iterative reasoning trajectory. Building on InftyThink’s paradigm of model-controlled iteration boundaries and explicit summarization, our approach proceeds in two stages: a cold-start stage that uses supervised fine-tuning to establish the basic iterative reasoning format, followed by an RL stage that optimizes strategic decisions through trajectory-level learning. We carefully design the rollout strategy, reward formulation, and policy gradient estimation tailored to InftyThink’s single-trajectory, multi-inference structure. This design separates format acquisition from strategy optimization, enabling the model to learn not only how to produce iterative reasoning, but also when to summarize, what to preserve, and how to effectively leverage self-generated summaries across iterations.
If you find our work helpful, feel free to give us a cite.
@misc{yan2026inftythinkplus,
title={InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning},
author={Yuchen Yan and Liang Jiang and Jin Jiang and Shuaicheng Li and Zujie Wen and Zhiqiang Zhang and Jun Zhou and Jian Shao and Yueting Zhuang and Yongliang Shen},
year={2026},
eprint={2602.06960},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.06960},
}
If you have any questions, please contact us by email: yanyuchen@zju.edu.cn