Downloads · 30 days
6
0% of all-time downloads
chreh/persuasive_language_detector
persuasive_language_detector is a text classification model from chreh. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
Given a sentence, our model predicts whether or not the sentence contains "persuasive" language, or language designed to elicit emotions or change readers' opinions. The model was tuned on the SemEval 2020 Task 11 dat…
Downloads · 30 days
6
0% of all-time downloads
All-time downloads
1.7K
Public
Parameters
334M
2.4 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
Given a sentence, our model predicts whether or not the sentence contains "persuasive" language, or language designed to elicit emotions or change readers' opinions. The model was tuned on the SemEval 2020 Task 11 dataset. However, we preprocessed the dataset to adapt it from multilabel technique classification and span-classification to our binary classification task.
There are two revisions:
bert-large-cased on our main branchxlm-roberta-base on our roberta branch.This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
Use the code below to get started with the model.
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-large-cased")
model = AutoModelForSequenceClassification.from_pretrained("chreh/persuasive_language_detector")
roberta branch (XLM RoBERTa)from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("xlm-roberta-base")
model = AutoModelForSequenceClassification.from_pretrained("chreh/persuasive_language_detector", revision="roberta")
Training data can be downloaded from the Semeval website.
The training was done using Huggingface Trainer on both our local machines and Intel Developer Cloud kernels, enabling us to prototype multiple models simultaneously.
All sentences containing spans of persuasive language techniques were labeled as persuasive language examples, while all others were labeled as examples of non-persuasive language.
The test data is from the test data of sem_eval_2020_task_11, which can be downloaded from the original website.
The test data contains 38.25% persuasive examples and non-persuasive examples 61.75%. Metrics can be found in the following section
Metrics are reported in the format (main_branch), (roberta branch)
Overall, the roberta branch performs better, and with faster inference times. Thus, we recommend users download from the roberta revision.