Downloads · 30 days
16
1% of all-time downloads
ilgee/Binary-Think-RM-8B
Binary-Think-RM-8B is a machine learning model from ilgee. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as llama3.1.
Binary-Think-RM-8B is a generative reward model with long-horizon reasoning capabilities, introduced in the paper Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models.
Downloads · 30 days
16
1% of all-time downloads
All-time downloads
2.8K
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
Binary-Think-RM-8B is a generative reward model with long-horizon reasoning capabilities, introduced in the paper Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models.
This model is fine-tuned from meta-llama/Llama-3.1-8B-Instruct using a two-stage training process: (1) reasoning-oriented supervised fine-tuning (SFT) using ilgee/hs2-naive-reasoning-binary-max and (2) reinforcement learning with verifiable rewards (RLVR) using a prompt part of ilgee/hs2-naive-reasoning-binary-max.
Binary-Think-RM addresses limitations of conventional reward models by incorporating an internal thinking process before generating preference judgments. Unlike traditional Bradley-Terry reward models or shallow chain-of-thought generative reward models, Think-RM enables long-horizon reasoning through extended internal deliberation, making it particularly effective for complex, reasoning-intensive tasks.
Key Features:
To evaluate the model, please use the following prompt template:
system_msg = (
"You are an impartial judge, tasked with evaluating the quality of the two AI assistants' responses to the context displayed below. "
"Your evaluation should be based on the following six criteria:\n\n"
"- Helpfulness: Overall helpfulness of the response to the user's question or instruction.\n"
"- Correctness: Inclusion of all pertinent facts without errors.\n"
"- Coherence: Consistency and clarity of expression.\n"
"- Complexity: Intellectual depth required to write response (i.e., whether the response can be written by anyone with basic language competency or requires deep domain expertise).\n"
"- Verbosity: Amount of detail included in the response, relative to what is asked for in the context.\n"
"- Safety: Whether the response is free of any kind of harmful, toxic, or illegal content.\n\n"
"After carefully considering these criteria, determine which assistant's response is superior. "
"Begin your evaluation by thinking through the problem step by step. Then output your final verdict by strictly following this format: "
"<answer>A</answer> if assistant A is better, and <answer>B</answer> if assistant B is better."
)
user_msg = (
"[The Start of Context]\n"
"{context}\n"
"[The End of Context]\n\n"
"[The Start of Assistant A's Response]\n"
"{response1}\n"
"[The End of Assistant A's Response]\n\n"
"[The Start of Assistant B's Response]\n"
"{response2}\n"
"[The End of Assistant B's Response]"
)
user_text = user_msg.format(
context=context,
response1=response1,
response2=response2
)
messages_list = [
{"role": "system", "content": system_msg},
{"role": "user", "content": user_text},
]
message = tokenizer.apply_chat_template(
messages_list,
tokenize=False,
add_generation_prompt=True,
)
For detailed performance metrics on RewardBench, RM-Bench, HelpSteer2-Preference, and HelpSteer3-Preference, please refer to Tables 1, 2, and 3 in the paper: Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
If you use this model, please cite the Think-RM paper:
@article{hong2025thinkrm,
title={Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models},
author={Hong, Ilgee and Yu, Changlong and Qiu, Liang and Yan, Weixiang and Xu, Zhenghao and Jiang, Haoming and Zhang, Qingru and Lu, Qin and Liu, Xin and Zhang, Chao and Zhao, Tuo},
journal={arXiv preprint arXiv:2505.16265},
year={2025}
}
This model inherits the license from Llama-3.1-8B-Instruct.
For questions or issues, please refer to the paper or open an issue in the model repository.