Downloads · 30 days
20
3% of all-time downloads
Jaew00Lee/Qwen3-4B-PRInTS
Qwen3-4B-PRInTS is a text generation model from Jaew00Lee. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This repository hosts the PRInTS (Process Reward via Information gain scoring and Trajectory Summarization) Qwen3-4B model. PRInTS is a generative process reward model for long-horizon information-seeking tasks.
Downloads · 30 days
20
3% of all-time downloads
All-time downloads
590
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
This repository hosts the PRInTS (Process Reward via Information gain scoring and Trajectory Summarization) Qwen3-4B model. PRInTS is a generative process reward model for long-horizon information-seeking tasks.
PRInTS (Process Reward via Information gain scoring and Trajectory Summary) is a generative PRM jointly trained with two key abilities for fine-grained guidance under the challenge of context accumulation.
Key Highlights:
Qwen3ForCausalLM, fine-tuned Large Language Model
The PRInTS (Qwen3-4B) model provides fine-grained guidance for information-seeking agents at test time, estimating step-level information-gain scores across n rollouts of the agents.
If you find this work useful, please consider citing us:
@article{lee2024prints,
title={PRInTS: Reward Modeling for Long-Horizon Information Seeking},
author={Jaewoo Lee and Archiki Prasad and Justin Chih-Yao Chen and Zaid Khan and Elias Stengel-Eskin and Mohit Bansal},
year={2025},
journal={arXiv preprint arXiv:2511.19314},
url={https://arxiv.org/abs/2511.19314},
}