Downloads · 30 days
5
31% of all-time downloads
gefero/conflict_detection_ROBERTA_based
conflict_detection_ROBERTA_based is a text classification model from gefero. Use it when you need a label for a piece of text. It is set up for transformers.
A fine-tuned XLM-RoBERTa model for detecting social conflict mentions in Spanish news articles. The model is trained on the "Conflicto Social en Noticias" dataset and achieves 91.07% macro-F1 score on test data, makin…
Downloads · 30 days
5
31% of all-time downloads
All-time downloads
16
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
A fine-tuned XLM-RoBERTa model for detecting social conflict mentions in Spanish news articles. The model is trained on the "Conflicto Social en Noticias" dataset and achieves 91.07% macro-F1 score on test data, making it suitable for automated content classification and conflict-related news filtering.
This is a binary text classification model based on FacebookAI/xlm-roberta-base fine-tuned to detect whether Spanish news articles discuss social conflict or not. The model was trained using a rigorous multi-seed approach (10 random seeds) to ensure robustness and generalization.
The classification task is binary:
The model achieved strong performance across multiple evaluation runs, with consistent metrics indicating reliable predictions on unseen test data.
This model can be used for:
This model can be integrated into:
This model is not suitable for:
Language: Model trained exclusively on Spanish news articles. Performance on other languages or dialects is unknown.
Domain: Model trained on news articles. Performance on other text types (social media, academic text, etc.) may be degraded.
Temporal bias: Dataset represents a specific time period. Linguistic evolution and emerging conflict narratives may not be captured.
Class balance: Dataset contains both conflict and non-conflict examples. Performance may vary based on class distribution in your specific use case.
Truncation: Text is truncated to 256 tokens (matching model's training setup). Very long articles may lose important context.
Context sensitivity: "Conflict" detection is based on textual patterns. Sarcasm, irony, or indirect references may be misclassified.
pip install transformers torch
from transformers import pipeline
# Initialize the model
classifier = pipeline(
"text-classification",
model="gefero/conflict_detection_ROBERTA_based"
)
# Classify text
texts = [
"El gobierno anunció nuevas políticas de seguridad social",
"Miles de personas protestaron en las calles contra las medidas económicas"
]
results = classifier(texts)
for text, result in zip(texts, results):
print(f"Text: {text[:50]}...")
print(f"Label: {result['label']} (score: {result['score']:.4f})\n")
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "gefero/conflict_detection_ROBERTA_based"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Prepare input
text = "Manifestantes se enfrentan con la policía"
inputs = tokenizer(
text,
truncation=True,
padding="max_length",
max_length=256,
return_tensors="pt"
)
# Get predictions
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
predictions = torch.argmax(logits, dim=-1)
confidence = torch.softmax(logits, dim=-1).max().item()
labels = {0: "NO_CONFLICTO", 1: "CONFLICTO"}
print(f"Prediction: {labels[predictions.item()]} (confidence: {confidence:.4f})")
Macro-F1 (primary metric): Average F1-score across both classes
Accuracy: Overall correctness of predictions
F1-CONFLICTO: F1-score specifically for the CONFLICTO class (conflict-related news)
Precision: True positives / (true positives + false positives)
Recall: True positives / (true positives + false negatives)
| Metric | Mean | Std Dev | Min | Max |
|---|---|---|---|---|
| Test Macro-F1 | 0.9107 | 0.0071 | 0.8999 | 0.9204 |
| Test Accuracy | 0.9394 | 0.0053 | 0.9319 | 0.9468 |
| Test F1-CONFLICTO | 0.8602 | 0.0108 | 0.8433 | 0.8746 |
| Test Precision | 0.8807 | 0.0296 | 0.8307 | 0.9182 |
| Test Recall | 0.8419 | 0.0247 | 0.8156 | 0.8883 |
| Seed | Dev Macro-F1 | Test Macro-F1 | Test Accuracy | Test F1-CONFLICTO | Test Precision | Test Recall |
|---|---|---|---|---|---|---|
| 0 | 0.9113 | 0.9007 | 0.9332 | 0.8439 | 0.8743 | 0.8156 |
| 1 | 0.9116 | 0.8999 | 0.9319 | 0.8433 | 0.8605 | 0.8268 |
| 7 | 0.8975 | 0.9092 | 0.9381 | 0.8580 | 0.8728 | 0.8436 |
| 13 | 0.9187 | 0.9140 | 0.9431 | 0.8639 | 0.9182 | 0.8156 |
| 42 | 0.9154 | 0.9127 | 0.9418 | 0.8622 | 0.9074 | 0.8212 |
| 100 | 0.9219 | 0.9136 | 0.9394 | 0.8665 | 0.8457 | 0.8883 |
| 123 | 0.9146 | 0.9194 | 0.9455 | 0.8736 | 0.8994 | 0.8492 |
| 2024 | 0.9098 | 0.9125 | 0.9406 | 0.8629 | 0.8830 | 0.8436 |
| 31337 | 0.9180 | 0.9204 | 0.9468 | 0.8746 | 0.9146 | 0.8380 |
| 65535 | 0.9168 | 0.9050 | 0.9332 | 0.8533 | 0.8307 | 0.8771 |
The model demonstrates excellent performance with:
Estimated CO2 emissions for multi-seed training approach: Low to minimal (Colab's data centers use renewable energy sources). Individual training runs are short (~10 min each) and performed on highly optimized infrastructure.
For detailed calculations, see ML Impact Calculator.
Architecture: Transformer-based sequence classification
Objective: Binary cross-entropy loss (standard for text classification)
Multilingual base: XLM-RoBERTa trained on 100+ languages, fine-tuned here for Spanish-specific conflict detection
If you use this model in research, please cite:
BibTeX:
@software{rosati2024conflictdetection,
author = {Rosati, Germán},
title = {XLM-RoBERTa Spanish Conflict Detection Classifier},
year = {2024},
publisher = {Hugging Face Hub},
url = {https://huggingface.co/gefero/conflict_detection_ROBERTA_based}
}
APA:
Rosati, G. (2024). XLM-RoBERTa Spanish Conflict Detection Classifier [Machine learning model]. Hugging Face Hub. Retrieved from https://huggingface.co/gefero/conflict_detection_ROBERTA_based