Downloads · 30 days
9
60% of all-time downloads
BaoNhan/cafebert-ViClickbait-2025
cafebert-ViClickbait-2025 is a text classification model from BaoNhan. Use it when you need a label for a piece of text. It is set up for transformers.
Fine-tuned from uitnlp/CafeBERT for binary Vietnamese clickbait detection.
Downloads · 30 days
9
60% of all-time downloads
All-time downloads
15
Public
Parameters
560M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
Fine-tuned from uitnlp/CafeBERT for binary Vietnamese clickbait detection.
StratifiedGroupKFold with seed 42.True.| Metric | Mean ± sample std |
|---|---|
| Test Macro-F1 | 0.8047 ± 0.0060 |
| Test accuracy | 0.8236 ± 0.0074 |
| Dev Macro-F1 | 0.8256 ± 0.0053 |
Representative seed: 22, selected only by development Macro-F1.
| seed | dev_macro_f1 | test_macro_f1 | test_accuracy |
|---|---|---|---|
| 22 | 0.8312 | 0.8043 | 0.8246 |
| 42 | 0.8249 | 0.7989 | 0.8158 |
| 202 | 0.8207 | 0.8108 | 0.8304 |
0: non-clickbait1: clickbaitfrom transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "BaoNhan/cafebert-ViClickbait-2025"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
title = "Tiêu đề bài báo"
lead = "Đoạn dẫn của bài báo"
inputs = tokenizer(title, lead, return_tensors="pt", truncation=True, max_length=256)
prediction = model(**inputs).logits.argmax(dim=-1).item()
print(model.config.id2label[prediction])
The dataset is small, temporally bounded to 2023–2025, and collected from eight Vietnamese news platforms. Results may not transfer to social media, other publishers, or emerging clickbait styles.