Downloads · 30 days
6
1% of all-time downloads
Himanshu167/AI-Response-Comparer
AI-Response-Comparer is a machine learning model from Himanshu167. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
This model is a fine-tuned version of microsoft/deberta-v3-base, optimized for Preference Classification (Reward Modeling). Instead of standard text classification, this model is designed to compare two AI-generated r…
Downloads · 30 days
6
1% of all-time downloads
All-time downloads
702
Public
Parameters
184M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors368 MB · 98%
How the weights are stored.
F16184M · 100%
From the Hugging Face model README
This model is a fine-tuned version of microsoft/deberta-v3-base, optimized for Preference Classification (Reward Modeling). Instead of standard text classification, this model is designed to compare two AI-generated responses to the same prompt and predict which one is higher quality or more "preferred."
The model is evaluated using the following criteria, comparing the predicted probability distribution [P(A), P(B), P(Tie)] against the ground truth:
Multi-class Log Loss (Primary):
Definition: Measures the distance between the predicted probability distribution and the actual labels. $$ L = -\frac{1}{N} \sum_{i=1}^{N} \sum_{j=1}^{M} y_{i,j} \log(p_{i,j}) $$
Variables: Where \(M = 3\) (representing Response A, Response B, and Tie).
Why: It rewards the model for assigning higher probabilities to the correct outcome and heavily penalizes high-confidence incorrect predictions.
Accuracy (Secondary):
Correct Predictions / Total Samples.The following results were achieved during final evaluation. Note that Accuracy was calculated using a local train/test split, while Log Loss follows the competition's evaluation framework.
| Metric | Value | Source/Split |
|---|---|---|
| Multi-class Log Loss | 1.0346 | Kaggle Competition Metric |
| Accuracy | 48.94% | Local Train/Test Split |
Note on Performance:
- Log Loss: This score reflects the model's ability to provide well-calibrated probabilities for the three classes (A, B, and Tie) as required by the Kaggle competition.
- Accuracy: This was monitored locally to ensure the model was successfully learning the preference patterns beyond a random baseline (33.33%).