Downloads · 30 days
10
8% of all-time downloads
hugsanaa/HatespeechLLM
HatespeechLLM is a machine learning model from hugsanaa. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
This repository hosts a fine-tuned version of the Mistral 7B (mistralai/Mistral-7B-Instruct-v0.3) language model for hate speech detection. The base model has been fine-tuned on a curated dataset containing various fo…
Downloads · 30 days
10
8% of all-time downloads
All-time downloads
119
Public
Repo size
336 MB
Likes
0
Public
Click a slice to open those files.
.safetensors336 MB · 99%
From the Hugging Face model README
This repository hosts a fine-tuned version of the Mistral 7B (mistralai/Mistral-7B-Instruct-v0.3) language model for hate speech detection. The base model has been fine-tuned on a curated dataset containing various forms of toxic, offensive, and hateful language across online platforms to make it suitable for detecting and classifying hate speech.
The model was fine-tuned using a binary labeled dataset of online texts including:
This model can be used to detect hate speech content online. It can be also used to be fine-tuned on more hate speech dataset.
If the data used for testing the model is collected from social media it is better to clean it by removing URLs, hashtags, mentions, and emojis.
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline
model_name = "hugsanaa/HatespeechLLM"
model = AutoModelForCausalLM.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name,
trust_remote_code=True,
max_length=512,
padding_side="left",
add_eos_token=True,
)
tokenizer.pad_token = tokenizer.eos_token
pipe = pipeline(task="text-generation",
model=model,
tokenizer=tokenizer,
max_new_tokens=10,
temperature=0.0
)
text = "generally women are forthright about reality and about everything else"
prompt = f"""
[INST] You are an AI model fine-tuned to detect hate speech. Below is a text, and you are required to determine whether it is hateful or non-hateful. Provide your answer as 'hateful' or 'non-hateful'. [/INST]
Text: {text}
Answer:
"""
result = pipe(prompt, pad_token_id=pipe.tokenizer.eos_token_id)
answer = result[0]['generated_text'].lower()
answer_index = answer.find("answer:") + len("answer:")
extracted_text = answer[answer_index:].strip()
if "non-" in extracted_text:
print("The sample provides non-hate speech")
elif "hate" in extracted_text:
print("The sample provides hate speech")
else:
print("Unable to detect whether the text belongs to hate or non-hate speech")