Downloads · 30 days
7
3% of all-time downloads
Mridul2003/identity-hate-detector
identity-hate-detector is a machine learning model from Mridul2003. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
7
3% of all-time downloads
All-time downloads
262
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
Use Model
from transformers import pipeline, AutoTokenizer, AutoModelForSequenceClassification
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
identity_model = AutoModelForSequenceClassification.from_pretrained("Mridul2003/identity-hate-detector").to(device)
identity_tokenizer = AutoTokenizer.from_pretrained("Mridul2003/identity-hate-detector")
identity_inputs = identity_tokenizer(final_text, return_tensors="pt", padding=True, truncation=True)
if 'token_type_ids' in identity_inputs:
del identity_inputs['token_type_ids']
identity_inputs = {k: v.to(device) for k, v in identity_inputs.items()}
with torch.no_grad():
identity_outputs = identity_model(**identity_inputs)
identity_probs = torch.sigmoid(identity_outputs.logits)
identity_prob = identity_probs[0][1].item()
not_identity_prob = identity_probs[0][0].item()
results["identity_hate_custom"] = identity_prob
results["not_identity_hate_custom"] = not_identity_prob
This repository contains a fine-tuned version of the unitary/toxic-bert model for binary classification of offensive language (labels: Offensive vs Not Offensive). The model has been specifically fine-tuned on a custom dataset due to limitations observed in the base model's performance — particularly with identity_hate related content.
unitary/toxic-bert)The original unitary/toxic-bert model is trained for multi-label toxicity detection with 6 categories:
While it performs reasonably well on generic toxicity, it struggles with edge cases involving identity-based hate speech — often:
We fine-tuned the model on a custom annotated dataset with two clear labels:
0: Not Identity Hate1: Identity HateThe new model simplifies the task into a binary classification problem, allowing more focused training for real-world moderation scenarios.
identity_hate, obscene, insult, and more nuanced samplesunitary/toxic-bertTrainer APInum_labels=2)| Metric | Value |
|---|---|
| Accuracy | ~92% |
| Precision / Recall | Balanced |