Downloads · 30 days
19
0% of all-time downloads
fgaim/tiroberta-abusiveness-detection
tiroberta-abusiveness-detection is a text classification model from fgaim. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-4.0.
This model is a fine-tuned version of TiRoBERTa on the TiALD dataset.
Downloads · 30 days
19
0% of all-time downloads
All-time downloads
7K
Public
Parameters
125M
499 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors499 MB · 98%
From the Hugging Face model README
This model is a fine-tuned version of TiRoBERTa on the TiALD dataset.
Tigrinya Abusive Language Detection (TiALD) Dataset is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of 13,717 YouTube comments annotated for abusiveness, sentiment, and topic tasks. The dataset includes comments written in both the Ge’ez script and prevalent non-standard Latin transliterations to mirror real-world usage.
⚠️ The dataset contains explicit, obscene, and potentially hateful language. It should be used for research purposes only. ⚠️
This work accompanies the paper "A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings".
from transformers import pipeline
tiald_pipe = pipeline("text-classification", model="fgaim/tiroberta-abusiveness-detection")
tiald_pipe("<text-to-classify>")
This model achieves the following results on the evaluation set:
"abusiveness_metrics": {
"accuracy": 0.8666666666666667,
"macro_f1": 0.8666502037288554,
"macro_precision": 0.8668478260869565,
"macro_recall": 0.8666666666666667,
"weighted_f1": 0.8666502037288554,
"weighted_precision": 0.8668478260869565,
"weighted_recall": 0.8666666666666667
}
The following hyperparameters were used during training:
The TiALD dataset and models designed to support:
Researchers and developers should avoid using this dataset for direct moderation or enforcement tasks without human oversight.
This research received IRB approval (Ref: KH2022-133) and followed ethical data collection and annotation practices, including informed consent of annotators.
If you use this model or the TiALD dataset in your work, please cite:
@misc{gaim-etal-2025-tiald-benchmark,
title = {A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings},
author = {Fitsum Gaim and Hoyun Song and Huije Lee and Changgeon Ko and Eui Jun Hwang and Jong C. Park},
year = {2025},
eprint = {2505.12116},
archiveprefix = {arXiv},
primaryclass = {cs.CL},
url = {https://arxiv.org/abs/2505.12116}
}
This dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0).