Downloads · 30 days
52
6% of all-time downloads
youngggggg/ToxiGen-ConPrompt
ToxiGen-ConPrompt is a feature extraction model from youngggggg. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
ToxiGen-ConPrompt is a pre-trained language model for implicit hate speech detection. The model is pre-trained on a machine-generated dataset for implicit hate speech detection (i.e., ToxiGen) using our proposing pre-…
Downloads · 30 days
52
6% of all-time downloads
All-time downloads
862
Public
Repo size
1.1 GB
Likes
3
Public
Click a slice to open those files.
.bin534 MB · 100%
From the Hugging Face model README
ToxiGen-ConPrompt is a pre-trained language model for implicit hate speech detection. The model is pre-trained on a machine-generated dataset for implicit hate speech detection (i.e., ToxiGen) using our proposing pre-training approach (i.e., ConPrompt).
<!-- Provide a quick summary of what the model is/does. --> <!-- {{ model_summary | default("", true) }} -->Before pre-training, we found out that some private information such as URLs exists in the machine-generated statements in ToxiGen. We anonymize such private information before pre-training to prevent any harm to our society. You can refer to the anonymization code we used in preprocess_toxigen.ipynb and we strongly emphasize to anonymize private information before using machine-generated data for pre-training.
The pre-training source of ToxiGen-ConPrompt includes toxic statements. While we use such toxic statements on purpose to pre-train a better model for implicit hate speech detection, the pre-trained model needs careful handling. Here, we states some behavior that can lead to potential misuse so that our model is used for the social good rather than misued unintentionally or maliciously.
While these behavior can lead to social good e.g., constructing training data for hate speech classifiers, one can potentially misuse the behaviors.
We strongly emphasize the need for careful handling to prevent unintentional misuse and warn against malicious exploitation of such behaviors.