Downloads · 30 days
173
84% of all-time downloads
opsd-genrm/DR-GRPO-Qwen3-4B
DR-GRPO-Qwen3-4B is a text generation model from opsd-genrm. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
📄 Paper: Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Downloads · 30 days
173
84% of all-time downloads
All-time downloads
207
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
📄 Paper: Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Pairwise generative reward model / LLM Judge trained from
Qwen/Qwen3-4B-Instruct-2507
on opsd-genrm/dedup_filtered_HS3
— a deduplicated and filtered version of
nvidia/HelpSteer3 —
with DR. GRPO. Given a context
and two candidate responses, the model first identifies the evaluation
criteria that matter for the specific task, then compares the two
responses step by step against those criteria, and finally emits a verdict
in <verdict>A</verdict> or <verdict>B</verdict>.
Send a single user message of the form below (no system prompt). The judge
expects each turn of the dialog and each candidate response to be wrapped
with <user>...</user> and <assistant>...</assistant> tags.
{context} is the user-side input. For a single-turn query it is just
<user>\n…question…\n</user>; for a multi-turn conversation it is the full
alternating dialog (must alternate user → assistant → user → … and end on
a <user> turn). {response_a} and {response_b} are the two candidate
replies, each wrapped in a single <assistant> block.
You are an impartial judge tasked with determining which of two assistant responses is better for the given context.
Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.
[Start of Context]
{context}
[End of Context]
[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]
[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]
Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>.
[Start of Context]
<user>
What is the capital of France?
</user>
[End of Context]
[Start of Assistant A's Response]
<assistant>
The capital of France is Paris.
</assistant>
[End of Assistant A's Response]
[Start of Assistant B's Response]
<assistant>
Lyon.
</assistant>
[End of Assistant B's Response]
[Start of Context]
<user>
I'm planning a 3-day trip to Tokyo next month. Any recommendations?
</user>
<assistant>
Sure — what kind of activities are you interested in (food, history, nightlife, shopping)?
</assistant>
<user>
Mostly food and history.
</user>
[End of Context]
[Start of Assistant A's Response]
<assistant>
Day 1: Tsukiji outer market for breakfast …
</assistant>
[End of Assistant A's Response]
[Start of Assistant B's Response]
<assistant>
Just go to Shibuya and figure it out when you get there.
</assistant>
[End of Assistant B's Response]
import re
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "opsd-genrm/DR-GRPO-Qwen3-4B"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(
REPO, torch_dtype="bfloat16", device_map="auto"
)
PROMPT = """You are an impartial judge tasked with determining which of two assistant responses is better for the given context.
Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.
[Start of Context]
{context}
[End of Context]
[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]
[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]
Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>."""
def format_context(messages: list[dict]) -> str:
"""Wrap a (multi-turn) dialog as alternating <user>/<assistant> blocks.
`messages` is a list of {"role": "user"|"assistant", "content": str},
must alternate starting with "user", and must end on a "user" turn.
"""
parts = []
for i, m in enumerate(messages):
expected = "user" if i % 2 == 0 else "assistant"
assert m["role"] == expected, "roles must alternate user/assistant/..."
parts.append(f"<{m['role']}>\n{m['content'].strip()}\n</{m['role']}>")
return "\n\n".join(parts)
def format_response(text: str) -> str:
return f"<assistant>\n{text.strip()}\n</assistant>"
def parse_verdict(text: str) -> str | None:
"""Prefers the last <verdict>...</verdict> block: returns the last 'A' or
'B' word inside it. If the verdict tag is missing or malformed, falls back
to the last standalone 'A'/'B' anywhere in the generated text. Returns
None only when no A/B token appears at all.
"""
blocks = re.findall(r"<verdict>(.*?)</verdict>", text, re.DOTALL)
if blocks:
ab = re.findall(r"\b(A|B)\b", blocks[-1])
if ab:
return ab[-1]
ab = re.findall(r"\b(A|B)\b", text)
return ab[-1] if ab else None
def judge(context_messages: list[dict], response_a: str, response_b: str) -> str | None:
user_msg = PROMPT.format(
context=format_context(context_messages),
response_a=format_response(response_a),
response_b=format_response(response_b),
)
inputs = tok.apply_chat_template(
[{"role": "user", "content": user_msg}],
add_generation_prompt=True, return_tensors="pt",
).to(model.device)
out = model.generate(
inputs,
max_new_tokens=8192,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
)
text = tok.decode(out[0, inputs.shape[-1]:], skip_special_tokens=True)
return parse_verdict(text)
# Single-turn:
verdict = judge(
context_messages=[{"role": "user", "content": "What is the capital of France?"}],
response_a="The capital of France is Paris.",
response_b="Lyon.",
)
print(verdict) # -> "A"
Generation config: temperature=0.7, top_p=0.8, top_k=20,
max_tokens=8192.
If you use this model, please cite:
@article{hong2026training,
title = {Training LLM Judges from Language Feedback via Position-Selective Self-Distillation},
author = {Hong, Ilgee and Yu, Changlong and Xu, Zhenghao and Liu, Xin and Zhang, Yuwei and Lu, Qin and Yin, Bing and Zhao, Tuo},
journal = {arXiv preprint arXiv:2609.38792},
year = {2026}
}
Released under Apache 2.0, inheriting from the Qwen3-4B-Instruct-2507 license.