Downloads · 30 days
362
0% of all-time downloads
bardsai/twitter-sentiment-pl-base
twitter-sentiment-pl-base is a text classification model from bardsai. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-4.0.
Twitter Sentiment PL (base) is a Polish-language sentiment analysis model fine-tuned from allegro/herbert-base-cased on a Polish translation of the TweetEval dataset (Barbieri et al., 2020). It predicts one of three s…
Downloads · 30 days
362
0% of all-time downloads
All-time downloads
171K
Public
Parameters
124M
1.5 GB on disk
Likes
4
Public
Click a slice to open those files.
.bin498 MB · 50%
How the weights are stored.
F32124M · 100%
From the Hugging Face model README
Twitter Sentiment PL (base) is a Polish-language sentiment analysis model fine-tuned from allegro/herbert-base-cased on a Polish translation of the TweetEval dataset (Barbieri et al., 2020). It predicts one of three sentiment classes for short, tweet-style Polish text.
pl)positive, negative, neutralWith the pipeline API:
from transformers import pipeline
nlp = pipeline("sentiment-analysis", model="bardsai/twitter-sentiment-pl-base")
nlp("Nigdy przegrana nie sprawiła mi takiej radości. Szczęście i Opatrzność mają znaczenie Gratuluje @pzpn_pl")
# [{'label': 'positive', 'score': 0.9997233748435974}]
Or loading the model and tokenizer directly:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("bardsai/twitter-sentiment-pl-base")
model = AutoModelForSequenceClassification.from_pretrained("bardsai/twitter-sentiment-pl-base")
Evaluated on the held-out test split (translated TweetEval, sentiment task) on an RTX 3090.
| Metric | Value |
|---|---|
| F1 (macro) | 0.658 |
| Precision (macro) | 0.655 |
| Recall (macro) | 0.662 |
| Accuracy | 0.662 |
| Samples per second | 129.9 |
This model is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, inherited from the base model allegro/herbert-base-cased, which is also distributed under CC BY 4.0.
You are free to share and adapt the model, including for commercial use, provided you give appropriate credit to:
If you use this model, please cite HerBERT and TweetEval:
@inproceedings{mroczkowski-etal-2021-herbert,
title = "{H}er{BERT}: Efficiently Pretrained Transformer-based Language Model for {P}olish",
author = "Mroczkowski, Robert and Rybak, Piotr and Wr{\'o}blewska, Alina and Gawlik, Ireneusz",
booktitle = "Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing",
year = "2021",
publisher = "Association for Computational Linguistics",
pages = "1--10",
}
@inproceedings{barbieri-etal-2020-tweeteval,
title = "{T}weet{E}val: Unified Benchmark and Comparative Evaluation for Tweet Classification",
author = "Barbieri, Francesco and Camacho-Collados, Jose and Espinosa Anke, Luis and Neves, Leonardo",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020",
year = "2020",
publisher = "Association for Computational Linguistics",
pages = "1644--1650",
}
At bards.ai we focus on providing machine learning expertise to our partners, particularly in NLP, computer vision and time series analysis. Our team is based in Wrocław, Poland.
If you use our model we'd love to hear about it. For questions or collaboration, contact us at [email protected].