Downloads · 30 days
17
4% of all-time downloads
mbkim/LifeTox_Moderator_350M
LifeTox_Moderator_350M is a text classification model from mbkim. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
17
4% of all-time downloads
All-time downloads
393
Public
Repo size
2.8 GB
Likes
2
Public
Click a slice to open those files.
.bin1.4 GB · 100%
From the Hugging Face model README
Dataset Card for LifeTox
As large language models become increasingly integrated into daily life, detecting implicit toxicity across diverse contexts is crucial. To this end, we introduce LifeTox, a dataset designed for identifying implicit toxicity within a broad range of advice-seeking scenarios. Unlike existing safety datasets, LifeTox comprises diverse contexts derived from personal experiences through open-ended questions. Our experiments demonstrate that RoBERTa fine-tuned on LifeTox matches or surpasses the zero-shot performance of large language models in toxicity classification tasks. These results underscore the efficacy of LifeTox in addressing the complex challenges inherent in implicit toxicity.
LifeTox Moderator 350M
LifeTox Moderator 350M is based on RoBERTa-large (350M). We fine-tuned this pre-trained model on LifeTox dataset. To use our model as a generalized moderator or specific pipelines, please refer to the paper 'LifeTox: Unveiling Implicit Toxicity in Life advice'. LifeTox Moderator 350M is trained as a toxicity scorer; output score >0 is safe, and <0 is unsafe.
BibTeX:
@article{kim2023lifetox,
title={LifeTox: Unveiling Implicit Toxicity in Life Advice},
author={Kim, Minbeom and Koo, Jahyun and Lee, Hwanhee and Park, Joonsuk and Lee, Hwaran and Jung, Kyomin},
journal={arXiv preprint arXiv:2311.09585},
year={2023}
}