Downloads · 30 days
147
5% of all-time downloads
UBC-NLP/DetoxLLM-7B
DetoxLLM-7B is a text generation model from UBC-NLP. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
<p align="center" <br <img src="./detoxllm.png" style="width: 20vw; min-width: 50px;" / <br <p
Downloads · 30 days
147
5% of all-time downloads
All-time downloads
3.2K
Public
Parameters
6.7B
27 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors27 GB · 100%
From the Hugging Face model README
This model card corresponds to the DetoxLLM-7B detoxification model based on LLaMA-2. The model is finetuned with Chain-of-Thought (CoT) explanation.
Paper: DetoxLLM: A Framework for Detoxification with Explanations (EMNLP 2024 Main)
Authors: Md Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V.S. Lakshmanan
Dataset: Dataset used to train this model can be found here.
Summary description and brief definition of inputs and outputs.
DetoxLLM is the first comprehensive end-to-end detoxification framework trained on cross-platform pseudo-parallel corpus. DetoxLLM further introduces explanation to promote transparency and trustworthiness. The framework also demonstrates robustness against adversarial toxicity.
Below we share some code snippets on how to get quickly started with running the model. First make sure to pip install -U transformers accelerate bitsandbytes, then copy the snippet from the section that is relevant for your usecase.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
torch.float16# pip install accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype=torch.float16)
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
torch.bfloat16from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype=torch.bfloat16)
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
bitsandbytesfrom transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(load_in_8bit=True)
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, quantization_config=quantization_config)
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(load_in_4bit=True)
model_name = "UBC-NLP/DetoxLLM-7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, quantization_config=quantization_config)
prompt = "Rewrite the following toxic input into non-toxic version. Let's break the input down step by step to rewrite the non-toxic version. You should first think about the expanation of why the input text is toxic. Then generate the detoxic output. You must preserve the original meaning as much as possible.\nInput: "
input = "Those shithead should stop talking and get the f*ck out of this place"
input_text = prompt+input+"\n"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, do_sample=False)
print(tokenizer.decode(outputs[0]))
The model is trained on cross-platform pseudo-parallel detoxification corpus generated using ChatGPT.
These models have certain limitations that users should be aware of.
The intended use of DetoxLLM is for the detoxification tasks. We aim to help researchers to build an end-to-end complete detoxification framework. DetoxLLM can also be regarded as a promising baseline to develop more robust and effective detoxification frameworks.
The development of large language models (LLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
If you use DetoxLLM for your scientific publication, or if you find the resources in this repository useful, please cite our paper as follows:
@inproceedings{khondaker-etal-2024-detoxllm,
title = "{D}etox{LLM}: A Framework for Detoxification with Explanations",
author = "Khondaker, Md Tawkat Islam and
Abdul-Mageed, Muhammad and
Lakshmanan, Laks V. S.",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.1066",
pages = "19112--19139",
}