Downloads · 30 days
9
33% of all-time downloads
erikbranmarino/Mistral-PRCT
Mistral-PRCT is a text classification model from erikbranmarino. Use it when you need a label for a piece of text. It is set up for peft. The card lists the license as mit.
A LoRA fine-tuned Mistral-7B model for detecting Population Replacement Conspiracy Theory (PRCT) content across (at least) Portuguese Telegram and Italian news headlines.
Downloads · 30 days
9
33% of all-time downloads
All-time downloads
27
Public
Repo size
252 MB
Likes
0
Public
Click a slice to open those files.
.safetensors294 MB · 100%
From the Hugging Face model README
A LoRA fine-tuned Mistral-7B model for detecting Population Replacement Conspiracy Theory (PRCT) content across (at least) Portuguese Telegram and Italian news headlines.
Mistral-PRCT is a LoRA adapter fine-tuned on Portuguese Telegram messages for detecting Population Replacement Conspiracy Theories. The model demonstrates strong cross-domain generalization, achieving competitive performance on Italian news headlines despite being trained exclusively on informal social media discourse.
| Dataset | F1-Macro | F1-Binary | Accuracy |
|---|---|---|---|
| Telegram PT (in-domain) | 0.819 | 0.700 | 0.896 |
| News ITA (cross-domain) | 0.753 | 0.688 | 0.771 |
Mistral-PRCT is a LoRA (Low-Rank Adaptation) fine-tuned version of Mistral-7B-Instruct-v0.3, specifically adapted for detecting Population Replacement Conspiracy Theory (PRCT) content. The model was trained on Portuguese Telegram messages and demonstrates robust cross-domain transfer to Italian news headlines.
Population Replacement Conspiracy Theories are false narratives claiming deliberate orchestration of demographic substitution through immigration. Main variants include:
These narratives are linked to extremist violence (Christchurch 2019, Utøya 2011) and pose serious threats to democratic discourse.
0: Non-PRCT content1: PRCT content (supports/mentions replacement narratives)Important: Should be used as part of a broader content moderation strategy, not as sole decision-maker.
| Metric | Score |
|---|---|
| Accuracy | 0.896 |
| Precision (Macro) | 0.797 |
| Recall (Macro) | 0.848 |
| F1-Macro | 0.819 |
| F1-Binary | 0.700 |
| Inference Time | 4.62s/sample |
| Metric | Score |
|---|---|
| Accuracy | 0.771 |
| Precision (Macro) | 0.748 |
| Recall (Macro) | 0.786 |
| F1-Macro | 0.753 |
| F1-Binary | 0.688 |
| Inference Time | 4.50s/sample |
Key Finding: Training on informal Portuguese Telegram enhances detection of implicit PRCT framing in formal Italian news, demonstrating effective cross-domain transfer from social media to journalistic discourse.
### Basic Usage
```pythonfrom transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torchLoad base model and tokenizer
base_model_name = "mistralai/Mistral-7B-Instruct-v0.3"
model = AutoModelForCausalLM.from_pretrained(
base_model_name,
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(base_model_name)Load LoRA adapter
model = PeftModel.from_pretrained(model, "erikbranmarino/Mistral-PRCT")Prepare prompt
text = "Your Portuguese or Italian text here"
prompt = f"""Classify if the following text contains Population Replacement Conspiracy Theory (PRCT) content.Text: {text}Classification (YES/NO):"""Generate prediction
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=10,
temperature=0.0,
do_sample=False
)
prediction = tokenizer.decode(outputs[0], skip_special_tokens=True)print(prediction)
### Batch Processing Example
```pythondef classify_batch(texts, model, tokenizer, batch_size=8):
"""Classify multiple texts efficiently"""
predictions = []for i in range(0, len(texts), batch_size):
batch = texts[i:i+batch_size]
prompts = [f"Classify PRCT: {text}" for text in batch] inputs = tokenizer(prompts, return_tensors="pt", padding=True).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=10, temperature=0.0) for output in outputs:
pred = tokenizer.decode(output, skip_special_tokens=True)
predictions.append(pred)return predictions
## Bias and Ethical Considerations
### Known Biases
- **Platform bias**: Optimized for Telegram-style informal discourse
- **Language bias**: Primarily Portuguese, with cross-lingual transfer to Italian
- **Temporal bias**: Training data from 2020-2024 may not capture evolving narratives
### Ethical Use
- ⚠️ **Not for automated censorship**: Requires human review
- ✅ **Research purposes**: Understanding conspiracy theory propagation
- ✅ **Content flagging**: Assisting moderators, not replacing them
- ❌ **Surveillance**: Not intended for monitoring individuals
We advocate for freedom of speech and constitutional rights. This tool should support informed moderation, not suppress legitimate discourse.
## Citation (to appear)
```bibtex@inproceedings{marino2025prct,
title={Population Replacement Conspiracy Theories Detection on Telegram and News Headlines:
benchmarking LLMs and BERT models in Portuguese and Italian},
author={Marino, Erik Bran and Vieira, Renata},
booktitle={Proceedings of PROPOR 2026},
year={2026}
}
## Model Card Authors
Erik Bran Marino (Universidade de Évora, HYBRIDS Project)
## Contact
- **Email**: [email protected]
- **Project**: MSCA HYBRIDS (Grant Agreement No. 101073351)
- **Institution**: Universidade de Évora, Portugal
## License
MIT License - Free for research and educational purposes.
---
**Developed as part of the HYBRIDS Marie Skłodowska-Curie Actions project**