Downloads · 30 days
8
20% of all-time downloads
Kushal0532/unblur-model
unblur-model is a machine learning model from Kushal0532. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Fine-tuned from answerdotai/ModernBERT-base. One shared backbone, three classification heads: clickbait (2-class), political leaning (3-class), sentiment (3-class).
Downloads · 30 days
8
20% of all-time downloads
All-time downloads
41
Public
Parameters
149M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt601 MB · 50%
From the Hugging Face model README
Fine-tuned from answerdotai/ModernBERT-base. One shared backbone, three
classification heads: clickbait (2-class), political leaning (3-class),
sentiment (3-class).
config.json + model.safetensors — backbone weights (HuggingFace format)tokenizer.json + tokenizer files — tokenizertask_heads.pt — classification head weights onlymodel_full.pt — full checkpoint (backbone + heads)Copy this folder to backend/models/ and start the backend. Or point
MODEL_REPO_ID at a private HF Hub repo and it'll pull from there.
Numbers below are from backend/evaluate.py on a 30-example hand-labelled
set covering the corner cases (ambiguous leaning, mixed sentiment, emotive-
but-legit headlines). The old numbers in this file were training-time
validation accuracy on the per-task datasets — not comparable, and
misleadingly high.
| Task | Accuracy | Macro F1 |
|---|---|---|
| Clickbait | 83% | 0.75 |
| Political leaning | 70% | 0.64 |
| Sentiment | 57% | 0.51 |
Latency: ~63ms avg, ~130ms p95 (CPU).
Test set is tiny (n=30), so treat these as ballpark — roughly ±15 points per task. Real eval needs a few hundred held-out examples. That's on the list.
Sentiment (57%) — trained on the wrong domain. The head learned on
tweet_eval, which is tweets. News headlines aren't tweets. Any headline
with a strong word ("crash", "devastate", "threat") gets read as
negative, even when the story is neutral or good news. Most of the errors
are neutral/positive headlines called negative. Used this dataset as a placeholder.
Clickbait (83%) — conflates tone with clickbait. Every error is a false positive: real headlines that happen to be loud ("GOP Tax Cuts Devastate Working Families", "Democrats Slam Republican Plan") get flagged. The training set was Buzzfeed-style listicle bait vs. clean wire copy, so the model keys on emotional language instead of the actual curiosity-gap pattern.
Political leaning (70%) — collapses toward center, barely sees right.
Only 2 of 6 right-leaning examples classified right; the rest went
center. Labels come from AllSides source ratings (config.py), so the
head learned publisher style, not article content, and the training mix
leans left-heavy. Need to use a larger test set to understand model choices better.
cardiffnlp/twitter-roberta onto an actual news corpus (already wired
up in datasets_loader.distill_sentiment_labels), or use a headline
sentiment set. Add a proper "neutral" bias since most news is neutral.