Downloads · 30 days
36
17% of all-time downloads
THU-KEG/WildReward-4B
WildReward-4B is a text classification model from THU-KEG. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
WildReward is a reward model trained on in-the-wild human-LLM interactions from the WildChat dataset. Unlike conventional reward models that rely on expensive human-annotated preference pairs, WildReward extracts impl…
Downloads · 30 days
36
17% of all-time downloads
All-time downloads
217
Public
Parameters
4B
8.1 GB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
WildReward is a reward model trained on in-the-wild human-LLM interactions from the WildChat dataset. Unlike conventional reward models that rely on expensive human-annotated preference pairs, WildReward extracts implicit reward signals from real-world user feedback through an automated pipeline.
WildReward is trained using ordinal regression (CORAL-like approach) on the WildFB dataset, which contains 186k high-quality instances filtered and refined from WildChat. Each instance is labeled with 5 levels of user satisfaction (Rejection, Error Correction, Neutral Ambiguity, Positive Engagement, Satisfaction).
Key Features:
WildFB Dataset (186k instances)
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_name = "THU-KEG/WildReward-4B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
def build_text(query, response, history_str=""):
"""Format input text for reward model scoring."""
text = f"""
# Task Description
You are an expert conversation evaluator. Your task is to judge the **User's Satisfaction** with the Assistant's response based on the conversation context.
Please rate the response on a scale of 1 to 5 integers.
# Scoring Criteria
[1] CLEARLY NEGATIVE / REJECTION
[2] CORRECTION / ERROR POINTER (Negative)
[3] NEUTRAL
[4] POSITIVE ENGAGEMENT
[5] CLEAR SATISFACTION
# Input Data
## Context (History)
{history_str}
## User Query
{query}
## Assistant Response
{response}
# Output
Based on the criteria above, please output ONLY the integer score (1, 2, 3, 4, or 5).
"""
return text.strip()
# Prepare query and response
query = "Explain quantum computing in simple terms."
response = "Quantum computing uses quantum bits or 'qubits' that can exist in multiple states simultaneously, unlike classical bits..."
# Build formatted text
text = build_text(query, response)
# Tokenize
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=4096).to(model.device)
# Get reward score
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
# CORAL / Ordinal Regression (output shape: 1, K-1)
probs = torch.sigmoid(logits)
reward = 1 + torch.sum(probs).item()
print(f"Reward score: {reward:.2f} (scale: 1-5)")
Architecture:
WildReward achieves competitive results on standard reward model benchmarks while demonstrating superior calibration properties. When applied to Online DPO, it significantly improves performance in mathematical reasoning, instruction following, and creative writing tasks.
@misc{peng2026wildrewardlearningrewardmodels,
title={WildReward: Learning Reward Models from In-the-Wild Human Interactions},
author={Hao Peng and Yunjia Qi and Xiaozhi Wang and Zijun Yao and Lei Hou and Juanzi Li},
year={2026},
eprint={2602.08829},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.08829},
}
Apache License 2.0
Note: This model card provides a brief overview. For detailed documentation on data collection, training, and deployment, please visit the GitHub repository.