Downloads · 30 days
32
3% of all-time downloads
yazoniak/twitter-emotion-pl-classifier
twitter-emotion-pl-classifier is a text classification model from yazoniak. Use it when you need a label for a piece of text. The card lists the license as gpl-3.0.
This model is a fine-tuned version of PKOBP/polish-roberta-8k for multi-label emotion and sentiment classification in Polish. It was trained on the TwitterEmo-PL-Refined dataset.
Downloads · 30 days
32
3% of all-time downloads
All-time downloads
1.1K
Public
Parameters
443M
1.8 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.8 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of PKOBP/polish-roberta-8k for multi-label emotion and sentiment classification in Polish. It was trained on the TwitterEmo-PL-Refined dataset.
The model predicts 8 emotion and sentiment labels simultaneously:
radość (joy), wstręt (disgust), gniew (anger), przeczuwanie (anticipation)pozytywny (positive), negatywny (negative), neutralny (neutral)sarkazm (sarcasm)max_length, e.g., 256-1024)| Metric | Score |
|---|---|
| F1 Macro | 0.8500 |
| F1 Micro | 0.8900 |
| F1 Weighted | 0.8895 |
| Exact Match Accuracy | 0.5125 |
| Subset Accuracy | 0.8900 |
| Validation Loss | 0.2761 |
| Label | F1 Score | Coverage |
|---|---|---|
| negatywny (negative) | 0.8553 | 42.4% |
| neutralny (neutral) | 0.8172 | 41.0% |
| pozytywny (positive) | 0.7814 | 17.4% |
| gniew (anger) | 0.7693 | 25.8% |
| radość (joy) | 0.7476 | 11.9% |
| wstręt (disgust) | 0.7337 | 20.4% |
| przeczuwanie (anticipation) | 0.7220 | 21.6% |
| sarkazm (sarcasm) | 0.5337 | 16.0% |
The model was trained on TwitterEmo-PL-Refined, which contains:
negatywny: 15,231 samples (42.4%)neutralny: 14,720 samples (41.0%)gniew: 9,252 samples (25.8%)przeczuwanie: 7,776 samples (21.6%)wstręt: 7,337 samples (20.4%)pozytywny: 6,248 samples (17.4%)sarkazm: 5,756 samples (16.0%)radość: 4,283 samples (11.9%)Model: PKOBP/polish-roberta-8k
Training samples: 28,737 (80%)
Validation samples: 7,184 (20%)
Hyperparameters:
- Learning rate: 1e-5
- Batch size: 32 (train), 32 (eval)
- Epochs: 4
- Weight decay: 0.03
- Warmup ratio: 0.1
- Dropout rate: 0.2
- Max gradient norm: 1.0
- Optimizer: AdamW
- LR scheduler: Cosine with warmup
- Early stopping patience: 3
- Mixed precision: BF16
Training strategy:
- Save strategy: Every 200 steps
- Evaluation strategy: Every 200 steps
- Best model selection: F1 Macro
- Total training steps: 3,600
- Best checkpoint: 3,400
Training was conducted on single NVIDIA RTX 3090 GPU using a stratified 80/20 train-validation split with the following progression:

The model's predictions can be improved using temperature scaling and optimized thresholds. Calibration analysis shows:
Per-label temperature scaling reduces calibration error (Expected Calibration Error - ECE):
| Label | Temperature | ECE Before | ECE After | Improvement |
|---|---|---|---|---|
radość | 1.066 | 0.0163 | 0.0166 | -1.8% |
wstręt | 1.117 | 0.0211 | 0.0152 | +27.9% |
gniew | 1.186 | 0.0308 | 0.0194 | +37.0% |
przeczuwanie | 1.102 | 0.0228 | 0.0237 | -3.9% |
pozytywny | 1.181 | 0.0280 | 0.0293 | -4.6% |
negatywny | 1.437 | 0.0594 | 0.0345 | +41.9% |
neutralny | 1.472 | 0.0696 | 0.0390 | +44.0% |
sarkazm | 1.078 | 0.0202 | 0.0202 | 0.0% |
Key findings:
neutralny, negatywny, and gniew benefit most from temperature scalingradość, przeczuwanie, pozytywny) show minor degradationPer-label F1-optimized thresholds (vs. default 0.5):
| Label | Optimal Threshold | F1 @ Optimal | F1 @ 0.5 | Improvement |
|---|---|---|---|---|
neutralny | 0.330 | 0.8211 | 0.8110 | +1.00% |
sarkazm | 0.330 | 0.5766 | 0.5256 | +5.10% |
przeczuwanie | 0.410 | 0.7276 | 0.7187 | +0.89% |
gniew | 0.440 | 0.7692 | 0.7676 | +0.16% |
negatywny | 0.450 | 0.8516 | 0.8511 | +0.05% |
wstręt | 0.460 | 0.7477 | 0.7464 | +0.13% |
pozytywny | 0.510 | 0.7864 | 0.7859 | +0.04% |
radość | 0.560 | 0.7572 | 0.7558 | +0.14% |
Key findings:
sarkazm shows the largest improvement (+5.10%) with a lower threshold (0.33)neutralny also benefits significantly (+1.00%) from a lower threshold (0.33)The model repository includes:
model.safetensors - Use with default threshold (0.5)calibration_artifacts.json - Contains temperature parameters and optimal thresholds
Recommendation: For production use, apply both temperature scaling and optimized thresholds for best performance.
This repository contains:
model.safetensors - Fine-tuned RoBERTa modeltokenizer.json, tokenizer_config.json - Polish RoBERTa tokenizerconfig.json - Model configurationcalibration_artifacts.json - Temperature scaling parameters and optimal thresholdspredict.py - Basic inference (threshold: 0.5)predict_calibrated.py - Calibrated inference (recommended)training_plots, calibration_reliability_diagramsrequirements.txt - Python dependenciesLICENSE - Full GPL-3.0 license textpip install -r requirements.txt
Or install dependencies manually:
pip install transformers torch numpy
The model expects @mentions to be anonymized, as they were during training. Both inference scripts automatically replace all @username mentions with @anonymized_account to match the training data distribution.
Use the predict.py script for basic inference with default threshold (0.5):
# From Hugging Face (default) - mentions are automatically anonymized
python predict.py "Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp"
# Example with mentions
python predict.py "@zgp_intervillage Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp"
# Preprocessed internally: "@anonymized_account Uwielbiam czekać..."
# From local model
python predict.py "Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp" --model-path ./
# With custom threshold
python predict.py "Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp" --model-path ./ --threshold 0.3
Example Output:
Loading model from: yazoniak/twitter-emotion-pl-classifier
Input text: Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp
Assigned Labels:
----------------------------------------
radość
pozytywny
sarkazm
All Labels (with probabilities):
----------------------------------------
✓ radość : 0.9574
wstręt : 0.0566
gniew : 0.0516
przeczuwanie : 0.0347
✓ pozytywny : 0.9782
negatywny : 0.0602
neutralny : 0.0336
✓ sarkazm : 0.5404
Use the predict_calibrated.py script for calibrated inference with temperature scaling and optimized thresholds:
# From Hugging Face with calibration (requires calibration_artifacts.json)
python predict_calibrated.py "Uwielbiam czekać na peronie 3 godziny! Gratulacje dla #zgp"
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import numpy as np
import re
def preprocess_text(text):
"""Preprocess text to match training data format."""
# Anonymize @mentions (IMPORTANT for best performance)
text = re.sub(r'@\w+', '@anonymized_account', text)
return text
# Load model
model_name = "yazoniak/twitter-emotion-pl-classifier"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
model.eval()
# Get labels from model config
labels = [model.config.id2label[i] for i in range(model.config.num_labels)]
# Prepare input with preprocessing
text = "@jan_kowalski To jest wspaniały dzień!"
preprocessed_text = preprocess_text(text) # "@anonymized_account To jest wspaniały dzień!"
inputs = tokenizer(preprocessed_text, return_tensors="pt", truncation=True, max_length=8192)
# Inference
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
# Get probabilities
probabilities = torch.sigmoid(logits).squeeze().numpy()
# Apply threshold
threshold = 0.5
predictions = {
label: float(prob)
for label, prob in zip(labels, probabilities)
if prob > threshold
}
print(predictions)
# Output: {'radość': 0.8734, 'pozytywny': 0.9156}
The model outputs logits for each of the 8 labels. To get predictions:
Preprocessing required: The model expects @mentions to be anonymized as @anonymized_account (matching training data). The provided inference scripts handle this automatically, but custom implementations must include this preprocessing step for optimal performance.
Sarcasm detection: The model struggles with Polish sarcasm (F1: 0.53), which is inherently difficult to detect in text for BERT models without additional context.
Class imbalance: Performance varies with label frequency:
negatywny, neutralny) perform bestradość, sarkazm) show lower F1 scoresTwitter-specific: The model is optimized for tweet-length texts (up to 8,192 tokens) with informal language, hashtags, and mentions.
If you use this model in your research or applications, please cite:
@model{yazoniak2025twitteremotionpl,
title={Polish Twitter Emotion Classifier (RoBERTa-8k)},
author={yazoniak},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/yazoniak/twitter-emotion-pl-classifier}
}
Also cite the base model and dataset:
@dataset{yazoniak_twitteremo_pl_refined_2025,
title = {TwitterEmo-PL-Refined: Polish Twitter Emotions (8 labels, refined)},
author = {yazoniak},
year = {2025},
url = {https://huggingface.co/datasets/yazoniak/TwitterEmo-PL-Refined}
}
@inproceedings{bogdanowicz2023twitteremo,
title = {TwitterEmo: Annotating Emotions and Sentiment in Polish Twitter},
author = {Bogdanowicz, S. and Cwynar, H. and Zwierzchowska, A. and Klamra, C. and Kiera{\'s}, W. and Kobyli{\'n}ski, {\L}.},
booktitle = {Computational Science -- ICCS 2023},
series = {Lecture Notes in Computer Science},
volume = {14074},
publisher = {Springer, Cham},
year = {2023},
doi = {10.1007/978-3-031-36021-3_20}
}
This model is released under the GNU General Public License v3.0 (GPL-3.0), inherited from the training dataset.
License Chain:
The complete GPL-3.0 license text is available in the LICENSE file in this repository, or at: https://www.gnu.org/licenses/gpl-3.0.html
For questions, issues, or feedback about this model, please open an issue in the model repository or contact the author through Hugging Face.
Model Version: v1.0 Last Updated: 2025-10-10