Downloads · 30 days
8
0% of all-time downloads
knowhate/HateBERTimbau-youtube
HateBERTimbau-youtube is a text classification model from knowhate. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc.
--- <img align="left" width="140" height="140" src="https://ilga-portugal.pt/files/uploads/2023/06/logoHATEcorespage-0001-1024x539.jpg" <p style="text-align: center;" This is the model card for…
Downloads · 30 days
8
0% of all-time downloads
All-time downloads
53.7K
Public
Parameters
109M
436 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors436 MB · 100%
From the Hugging Face model README
HateBERTimbau-YouTube is a transformer-based encoder model for identifying Hate Speech in Portuguese social media text. It is a fine-tuned version of HateBERTimbau model, retrained on a dataset of 23,912 YouTube comments specifically focused on Hate Speech.
You can use this model directly with a pipeline for text classification:
from transformers import pipeline
classifier = pipeline('text-classification', model='knowhate/HateBERTimbau-youtube')
classifier("as pessoas tem que perceber que ser 'panasca' não é deixar de ser homem, é deixar de ser humano 😂😂")
[{'label': 'Hate Speech', 'score': 0.9228119850158691}]
Or this model can be used by fine-tuning it for a specific task/dataset:
from transformers import AutoTokenizer, AutoModelForSequenceClassification, TrainingArguments, Trainer
from datasets import load_dataset
tokenizer = AutoTokenizer.from_pretrained("knowhate/HateBERTimbau-youtube")
model = AutoModelForSequenceClassification.from_pretrained("knowhate/HateBERTimbau-youtube")
dataset = load_dataset("knowhate/youtube-train")
def tokenize_function(examples):
return tokenizer(examples["sentence1"], examples["sentence2"], padding="max_length", truncation=True)
tokenized_datasets = dataset.map(tokenize_function, batched=True)
training_args = TrainingArguments(output_dir="hatebertimbau", evaluation_strategy="epoch")
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized_datasets["train"],
eval_dataset=tokenized_datasets["validation"],
)
trainer.train()
23,912 YouTube comments associated with offensive content were used to fine-tune the base model.
The dataset used to test this model was: knowhate/youtube-test
| Dataset | Precision | Recall | F1-score |
|---|---|---|---|
| knowhate/youtube-test | 0.856 | 0.892 | 0.874 |
Currently in Peer Review
@article{
}
This work was funded in part by the European Union under Grant CERV-2021-EQUAL (101049306). However the views and opinions expressed are those of the author(s) only and do not necessarily reflect those of the European Union or Knowhate Project. Neither the European Union nor the Knowhate Project can be held responsible.