Downloads · 30 days
18
7% of all-time downloads
przvl/PopEuroBERT-binary-610m
PopEuroBERT-binary-610m is a text classification model from przvl. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
1. Overview 2. Usage 3. Training Data 4. Training Procedure 5. Evaluation 6. Limitations 7. Ethical Considerations 8. License 9. Citation
Downloads · 30 days
18
7% of all-time downloads
All-time downloads
242
Public
Parameters
609M
2.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors2.4 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of EuroBERT-610m on the PopBERT dataset (sentence-level annotated German Bundestag speeches) for populist rhetoric classification. It predicts whether a given speech excerpt contains populist language.
Key Features:
To use the model in Python:
import torch
from transformers import AutoTokenizer
from transformers import AutoModelForSequenceClassification
model_id = "przvl/PopEuroBERT-binary-610m"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, trust_remote_code=True
)
# define text to be predicted
text = (
"Aber Ihnen fehlt eben der Mut, Ihnen fehlen die Visionen, um sich"
"gegen die Konzerne und gegen die Lobbygruppen zur Wehr zu setzen."
)
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# get classification probability
logits = outputs.logits
probs = torch.softmax(logits, dim=-1) # shape [1, 2]
populist_prob = probs[0, 1].item() # probability of class=1 (populist)
# use decision threshold 0.43
threshold = 0.43
label = "Populist" if populist_prob > threshold else "Neutral"
print(f"Predicted class: {label} (Confidence: {populist_prob:.2f})")
Predicted class: Populist (Confidence: 0.72)
Use decision threshold 0.43 for balanced performance.
train/test: 7017/1758populist = 1, neutral = 0).256 tokens.| Parameter | Value |
|---|---|
| Learning Rate | 1e-05 |
| Weight Decay | 0.1 |
| Gradient Accumulation | 1 |
| Warmup Ratio | 0.1 |
| Epochs | 3 |
| Batch Size | 128 |
| Max Length | 256 |
For transparency, we compare this model with its smaller variant (PopEuroBERT-210m), both trained and evaluated on the same dataset and splits.
| Model | Accuracy | Precision | Recall | F1 Score | Loss |
|---|---|---|---|---|---|
| 210M | 75.99% | 73.78% | 80.66% | 77.07% | 0.4959 |
| 610M (this) | 80.26% | 78.42% | 83.50% | 80.89% | 0.4631 |
| Model | Threshold | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|---|
| 210M | 0.56 | 76.00% | 76.00% | 76.00% | 76.00% |
| 610M (this) | 0.43 | 79.81% | 76.63% | 85.78% | 80.94% |
PopEuroBERT-610m consistently outperforms the 210m variant across all metrics. It especially improves recall and F1 score, suggesting better identification of populist speech. The decision threshold (0.43) was tuned for balanced performance.
0.43) was optimized for this dataset but may need adjustment for other corpora.Released under the Apache 2.0 License.
If you use this model or its methodology, please cite:
The original EuroBERT paper:
@misc{boizard2025eurobertscalingmultilingualencoders,
title={EuroBERT: Scaling Multilingual Encoders for European Languages},
author={Nicolas Boizard and Hippolyte Gisserot-Boukhlef and Duarte M. Alves and André Martins and Ayoub Hammal and Caio Corro and Céline Hudelot and Emmanuel Malherbe and Etienne Malaboeuf and Fanny Jourdan and Gabriel Hautreux and João Alves and Kevin El-Haddad and Manuel Faysse and Maxime Peyrard and Nuno M. Guerreiro and Patrick Fernandes and Ricardo Rei and Pierre Colombo},
year={2025},
eprint={2503.05500},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2503.05500}
}
The PopBERT dataset source:
@article{Erhard_Hanke_Remer_Falenska_Heiberger_2025,
title={PopBERT. Detecting Populism and Its Host Ideologies in the German Bundestag},
volume={33},
DOI={10.1017/pan.2024.12},
number={1},
journal={Political Analysis},
author={Erhard, Lukas and Hanke, Sara and Remer, Uwe and Falenska, Agnieszka and Heiberger, Raphael Heiko},
year={2025},
pages={1–17}
}