Downloads · 30 days
18
19% of all-time downloads
THU-KEG/LongTraceRL-4B
LongTraceRL-4B is a reinforcement learning model from THU-KEG. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
[](https://arxiv.org/abs/2605.31584) [](https://github.com/THU-KEG/LongTraceRL)
Downloads · 30 days
18
19% of all-time downloads
All-time downloads
94
Public
Parameters
4B
8.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
LongTraceRL-4B is a 4-billion parameter reasoning model trained with reinforcement learning on long-context multi-hop QA tasks using trajectory-based tiered distractors and entity-level rubric rewards.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("THU-KEG/LongTraceRL-4B")
tokenizer = AutoTokenizer.from_pretrained("THU-KEG/LongTraceRL-4B")
@misc{lin2026longtracerllearninglongcontextreasoning,
title={LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards},
author={Nianyi Lin and Jiajie Zhang and Lei Hou and Juanzi Li},
year={2026},
eprint={2605.31584},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.31584},
}