Downloads · 30 days
10
18% of all-time downloads
pjait/deberta-disinfo-detection
deberta-disinfo-detection is a text classification model from pjait. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of microsoft/deberta-v3-base for binary disinformation detection. It classifies news articles as either credible (0) or disinformation (1).
Downloads · 30 days
10
18% of all-time downloads
All-time downloads
57
Public
Parameters
184M
738 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors738 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of microsoft/deberta-v3-base for binary disinformation detection. It classifies news articles as either credible (0) or disinformation (1).
| Parameter | Value |
|---|---|
| Learning rate | 1e-05 |
| Batch size (train) | 16 |
| Batch size (eval) | 16 |
| Epochs | 5 |
| Weight decay | 0.01 |
| Warmup ratio | 0.06 |
| FP16 | True |
| Max sequence length | 512 |
| Seed | 42 |
| Eval steps | 100 |
| Best model selection | binary_f1_pos |
| Metric | Value |
|---|---|
| Binary F1 (positive) | 0.9041 |
| Macro F1 | 0.9342 |
| Accuracy | 0.948 |
| AUC-ROC | 0.9864 |
| Precision | 0.9223 |
| Recall | 0.9485 |
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("pjait/deberta-v3-base-disinfo-task1-binary")
model = AutoModelForSequenceClassification.from_pretrained("pjait/deberta-v3-base-disinfo-task1-binary")
text = "Your article text here..."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
probability = torch.sigmoid(logits).item()
prediction = "disinformation" if probability >= 0.5 else "credible"
print(f"Prediction: {prediction} (probability: {probability:.4f})")
The model was trained on the about 5k articles dataset for Task 1 (binary classification), which contains news articles annotated by multiple annotators for credibility assessment.
If you use this model, please cite the paper (to do, currently paper under review).