Downloads · 30 days
0
narinzar/reward-model-from-scratch
reward-model-from-scratch is a text classification model from narinzar. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
A small reward model trained from scratch on pairwise preference comparisons: a scalar reward head on top of distilbert-base-uncased, trained with the Bradley-Terry loss (-logsigmoid(rchosen - rrejected)), then evalua…
Downloads · 30 days
0
Access
Public
Updated Jul 8, 2026
Repo size
265 MB
Likes
0
Public
Click a slice to open those files.
.pt265 MB · 100%
From the Hugging Face model README
A small reward model trained from scratch on pairwise preference comparisons: a
scalar reward head on top of distilbert-base-uncased, trained with the
Bradley-Terry loss (-log_sigmoid(r_chosen - r_rejected)), then evaluated by
held-out pairwise accuracy and a calibration analysis (reliability diagram + ECE).
Code: https://github.com/narinzar/reward-model-from-scratch
Reward modeling / text classification. Given two responses to the same prompt, the model assigns each a scalar reward; the preferred response should score higher.
distilbert-base-uncasedLinear(hidden -> 1) producing one
scalar reward per sequenceMeasured on a single RTX 5090 (sm_120, CUDA 12.8 wheels). 1447 synthetic preference pairs (1158 train / 289 held-out eval), 3 epochs, batch size 16, lr 2e-5.
| Metric | Value |
|---|---|
| Held-out pairwise accuracy | 1.0000 |
| Expected Calibration Error (ECE, 10 bins) | 0.0078 |
The perfect held-out accuracy reflects that the synthetic label rule (longer-and-more-polite) is fully learnable at this scale. On a real, noisier preference corpus expect accuracy below 1.0 and a larger calibration gap.
Synthetic. Preference pairs are generated and labelled by a transparent rule (longer-and-more-polite is preferred), clearly marked as synthetic in the source. This is a from-scratch demonstration, not a production reward model. Swap in a real preference corpus by replacing one data loader.
reward_model.pt — checkpoint dict with state_dict, base_model_name,
and max_length. Load with the RewardModel class from the GitHub repo.MIT