Downloads · 30 days
7
2% of all-time downloads
trustyai/tci_plus
tci_plus is a machine learning model from trustyai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a facebook/bart-large fine-tuned on non-toxic inputs from lmsys/toxic-chat dataset.
Downloads · 30 days
7
2% of all-time downloads
All-time downloads
366
Public
Parameters
406M
6.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt3.3 GB · 67%
From the Hugging Face model README
This model is a facebook/bart-large fine-tuned on non-toxic inputs from lmsys/toxic-chat dataset.
This model is not intended to be used for plain inference despite it is unlikely to generate toxic content. It is intended to be used instead as "utility model" for detecting and fixing toxic content as its token probability distributions will likely differ from comparable models not trained/fine-tuned over non-toxic data.
Its name tci_plus refers to the G+ model in Detoxifying Text with MaRCo: Controllable Revision with Experts and Anti-Experts.
It can be used within TrustyAI's TMaRCo tool for detoxifying text, see https://github.com/trustyai-explainability/trustyai-detoxify/.
This model is intended to be used as "utility model" for detecting and fixing toxic content as its token probability distributions will likely differ from comparable models not trained/fine-tuned over toxic data.
This model is fine-tuned over non-toxic inputs from the lmsys/toxic-chat dataset and it is very likely to produce toxic content. For this reason this model should only be used in combination with other models for the sake of detecting / fixing toxic content.
Use the code below to start using the model for text detoxification.
from trustyai.detoxify import TMaRCo
tmarco = TMaRCo(expert_weights=[-1, 3])
tmarco.load_models(["trustyai/tci_minus", "trustyai/tci_plus"])
tmarco.rephrase(["white men can't jump"])
This model has been trained on non-toxic inputs from the lmsys/toxic-chat dataset.
Training data from the lmsys/toxic-chat dataset.
This model has been fine tuned with the following code:
from trustyai.detoxify import TMaRCo
dataset_name = 'lmsys/toxic-chat'
data_dir = ''
perc = 100
td_columns = ['model_output', 'user_input', 'human_annotation', 'conv_id', 'jailbreaking', 'openai_moderation',
'toxicity']
target_feature = 'toxicity'
content_feature = 'user_input'
model_prefix = 'toxic_chat_input_'
tmarco.train_models(perc=perc, dataset_name=dataset_name, expert_feature=target_feature, model_prefix=model_prefix,
data_dir=data_dir, content_feature=content_feature, td_columns=td_columns)
This model has been trained with the following hyperparams:
training_args = TrainingArguments(
evaluation_strategy="epoch",
learning_rate=2e-5,
weight_decay=0.01
)
Test data from the lmsys/toxic-chat dataset.
The model was evaluated using perplexity metric.
Perplexity: 1.04