Downloads · 30 days
7
1% of all-time downloads
dk3156/toxic_tweets_model
toxic_tweets_model is a text classification model from dk3156. Use it when you need a label for a piece of text. It is set up for transformers.
Downloads · 30 days
7
1% of all-time downloads
All-time downloads
893
Public
Repo size
536 MB
Likes
0
Public
Click a slice to open those files.
.bin268 MB · 100%
From the Hugging Face model README
#distilbert-base-uncased
This model is based on the pre-trained model [distilbert-base-uncased] and was fine-tuned on a dataset of tweets from Kaggle's Toxic Comment Classification Challenge
The model has been trained on the toxicity of tweets ranging from toxic, severe toxic, obscene, threat, insult, hate speech
The model predicts 6 signals of toxicity:
Toxic Severe Toxic Obscene Threat Insult Hate Speech
A value between 0 and 1 is predicted for each signal.
The model was created to be used as a toxicity detector of tweets based on the six categories. Other forms of toxicity from tweets may not be calculated with this model.
The model can be used directly with a text-classification pipeline:
>>> from transformers import pipeline
>>> text = "Your vandalism to the Matt Shirvington article has been reverted. Please don't do it again, or you will be banned."
>>> pipe = pipeline("text-classification", model="dk3156/toxic_tweets_model")
>>> pipe(text, return_all_scores=True)
[[{'label0': 'score': 0.02},
{'label1': 'score': 0.0},
{'label2': 'score': 0.0},
{'label3': 'score': 0.0},
{'label4': 'score': 0.0},
{'label5': 'score': 0.0}]]
The pre-trained model was fine-tuned for sequence classification using the following hyperparameters, which were selected from a validation set:
The optimizer used was AdamW.