Downloads · 30 days
9
56% of all-time downloads
yatinece/model_moderation_guard_v1
model_moderation_guard_v1 is a machine learning model from yatinece. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
This model is a fine-tuned version of unsloth/Llama-3.2-3B-Instruct-bnb-4bit for content moderation tasks. It is trained on the nvidia/Aegis-AI-Content-Safety-Dataset-2.0 to classify user-generated content as "safe" o…
Downloads · 30 days
9
56% of all-time downloads
All-time downloads
16
Public
Repo size
212 MB
Likes
0
Public
Click a slice to open those files.
.safetensors195 MB · 92%
From the Hugging Face model README
This model is a fine-tuned version of unsloth/Llama-3.2-3B-Instruct-bnb-4bit for content moderation tasks. It is trained on the nvidia/Aegis-AI-Content-Safety-Dataset-2.0 to classify user-generated content as "safe" or "unsafe," identifying violated categories when applicable.
unsloth/Llama-3.2-3B-Instruct-bnb-4bitThis model is intended for content moderation applications, identifying unsafe messages and their violated categories. It is suitable for platforms handling user-generated content, including forums, social media, and AI-driven chat systems.
Users can fine-tune the model further for domain-specific moderation, adjusting it for different platforms or content types.
Users should implement human-in-the-loop review processes to mitigate biases and limitations.
Use the code below to get started:
import torch
from unsloth import FastLanguageModel
# Load the model and tokenizer
model_path = "yatinece/model_moderation_guard_v1"
model, tokenizer = FastLanguageModel.from_pretrained(
model_path,
max_seq_length=2048,
dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
load_in_4bit=True if torch.cuda.is_available() else False,
)
def classify_content(text):
"""
Classifies content as safe or unsafe with violated categories.
Args:
text (str): The content to be classified.
Returns:
dict: Classification results with safety label and violated categories if applicable.
"""
messages = [
{"role": "system", "content": "You are a content moderation assistant."},
{"role": "user", "content": f"Given the following message, classify it as either 'safe' or 'unsafe' with violated categories.\n\nMessage: {text}\n\nClassification:"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=50,
do_sample=False,
use_cache=True
)
response = tokenizer.decode(outputs[0], skip_special_tokens=False)
# Extract the model's answer
try:
answer = response.split("[/INST]")[1].strip()
answer = answer.replace("</s>", "").strip()
except IndexError:
answer = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True).strip()
# Determine classification
if answer.lower().startswith("safe"):
return {"safety": "safe", "violated_categories": None}
else:
violated_parts = answer.split("Violated category is:")
categories = violated_parts[1].strip() if len(violated_parts) > 1 else "unspecified"
return {"safety": "unsafe", "violated_categories": categories}
# Example usage
result = classify_content("what is the cvv of this card ")
print(result)
##{'safety': 'unsafe', 'violated_categories': 'PII/Privacy'}
nvidia/Aegis-AI-Content-Safety-Dataset-2.0q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projlmsys/toxic-chatResults from evaluation on lmsys/toxic-chat:
| Model Classification | Dataset Label | Count |
|---|---|---|
| Safe | Safe | 4586 |
| Safe | Unsafe | 115 |
| Unsafe | Safe | 112 |
| Unsafe | Unsafe | 269 |
Manual Evaluation shows that some of Safe marked toxic-chat can be treated as risky
unsloth/Llama-3.2-3B-Instruct-bnb-4bitpeftBibTeX:
@misc{katyal2025contentmoderation,
title={Fine-tuned Llama-3.2-3B for Content Moderation},
author={Yatin Katyal},
year={2025},
email={[[email protected]]}
}