Downloads · 30 days
32
12% of all-time downloads
olehmell/xlm-roberta-posts-manipulation-classifier
xlm-roberta-posts-manipulation-classifier is a text classification model from olehmell. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This model detects propaganda and manipulation techniques in Ukrainian and Russian text. It is a fine-tuned version of FacebookAI/xlm-roberta-large trained on a bilingual subset of the UNLP 2025 Shared Task dataset fo…
Downloads · 30 days
32
12% of all-time downloads
All-time downloads
262
Public
Parameters
278M
2.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
This model detects propaganda and manipulation techniques in Ukrainian and Russian text. It is a fine-tuned version of FacebookAI/xlm-roberta-large trained on a bilingual subset of the UNLP 2025 Shared Task dataset for multi-label classification of manipulation techniques.
Its multilingual architecture makes it effective at understanding nuances in both Ukrainian and Russian, including code-mixed contexts.
The model performs multi-label text classification, identifying 5 major manipulation categories. A single text can contain multiple techniques.
The model was trained on the dataset from the UNLP 2025 Shared Task on manipulation technique classification.
The model was fine-tuned using the following hyperparameters:
| Parameter | Value |
|---|---|
| Base Model | FacebookAI/xlm-roberta-large |
| Learning Rate | 2e-5 |
| Train Batch Size | 16 |
| Eval Batch Size | 32 |
| Epochs | 10 |
| Max Sequence Length | 512 |
| Optimizer | AdamW |
| Loss Function | BCEWithLogitsLoss (with class weights) |
First, install the necessary libraries:
pip install transformers torch sentencepiece
Here is how to use the model to classify a single piece of text:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
# Define model and label names
model_name = "olehmell/ukr-rus-manipulation-detector-xlm-roberta" # Hypothetical model name
labels = [
'emotional_manipulation',
'fear_appeals',
'bandwagon_effect',
'selective_truth',
'cliche'
]
# Load pretrained model and tokenizer
model = AutoModelForSequenceClassification.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Prepare text (can be Ukrainian or Russian)
text = "Все эксперты уже давно это подтвердили, только вы не понимаете, что происходит на самом деле."
# Tokenize and predict
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.sigmoid(outputs.logits)
# Get detected techniques
threshold = 0.5
detected_techniques = {}
for i, score in enumerate(predictions[0]):
if score > threshold:
detected_techniques[labels[i]] = f"{score:.2f}"
if detected_techniques:
print("Detected techniques:")
for technique, score in detected_techniques.items():
print(f"- {technique} (Score: {score})")
else:
print("No manipulation techniques detected.")
The model achieves the following performance on the evaluation set:
| Metric | Value |
|---|---|
| F1 Macro | 0.44 |
| F1 Micro | TBD |
| Hamming Loss | TBD |
If you use this model in your research, please cite the following:
@misc{ukrainian-russian-manipulation-xlm-roberta-2025,
author = {Oleh Mell},
title = {Ukrainian/Russian Manipulation Detector - XLM-RoBERTa},
year = {2025},
publisher = {Hugging Face},
url = {[https://huggingface.co/olehmell/ukr-rus-manipulation-detector-xlm-roberta](https://huggingface.co/olehmell/ukr-rus-manipulation-detector-xlm-roberta)}
}
@inproceedings{unlp2025shared,
title={UNLP 2025 Shared Task on Techniques Classification},
author={UNLP Workshop Organizers},
booktitle={UNLP 2025 Workshop},
year={2025},
url={[https://github.com/unlp-workshop/unlp-2025-shared-task](https://github.com/unlp-workshop/unlp-2025-shared-task)}
}
This model is licensed under the Apache 2.0 License.