Downloads · 30 days
16
2% of all-time downloads
TheMrguiller/ToxDialogDefender
ToxDialogDefender is a text classification model from TheMrguiller. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This model is part of the research presented in "Mitigating Toxicity in Dialogue Agents through Adversarial Reinforcement Learning," a conference paper addressing dialog agent toxicity by mitigating it at three levels…
Downloads · 30 days
16
2% of all-time downloads
All-time downloads
798
Public
Parameters
139M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin557 MB · 48%
How the weights are stored.
F32139M · 100%
From the Hugging Face model README
This model is part of the research presented in "Mitigating Toxicity in Dialogue Agents through Adversarial Reinforcement Learning," a conference paper addressing dialog agent toxicity by mitigating it at three levels: explicit, implicit, and contextual. It is a model capable of predicting toxicity given a history and a response to it. It is designed for dialog agents. To use it correctly, please follow the schematics below:
[HST]Hi, how are you?[END]I am doing fine[ANS]I hope you die.
The token [HST] initiates the history of the conversation, and each turn pair is separated by [END]. The token [ANS] indicates the start of the response to the last utterance. I will update this card, but right now, I am developing a bigger project with these, so I do not have the time to indicate all the results.
The datasets used to train the model were the Dialogue Safety dataset and Bot Adversarial dataset.