Downloads · 30 days
36
0% of all-time downloads
holistic-ai/rejection_detection
rejection_detection is a text classification model from holistic-ai. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This model was originally developed and fine-tuned by Protect AI. It is a fine-tuned version of distilroberta-base, trained on multiple datasets containing rejection responses from LLMs and standard outputs from RLHF…
Downloads · 30 days
36
0% of all-time downloads
All-time downloads
137K
Public
Parameters
82.1M
657 MB on disk
Likes
1
Public
Click a slice to open those files.
.onnx329 MB · 50%
From the Hugging Face model README
This model was originally developed and fine-tuned by Protect AI. It is a fine-tuned version of distilroberta-base, trained on multiple datasets containing rejection responses from LLMs and standard outputs from RLHF datasets.
The goal of this model is to detect LLM rejections when a prompt does not pass content moderation. It classifies responses into two categories:
0: Normal output1: Rejection detectedOn the evaluation set, the model achieves:
The model is designed to identify rejection responses in LLM outputs, particularly where a refusal or safeguard message is generated.
Limitations:
distilroberta-base, it is case-sensitive.from transformers import AutoTokenizer, AutoModelForSequenceClassification, pipeline
import torch
tokenizer = AutoTokenizer.from_pretrained("ProtectAI/distilroberta-base-rejection-v1")
model = AutoModelForSequenceClassification.from_pretrained("ProtectAI/distilroberta-base-rejection-v1")
classifier = pipeline(
"text-classification",
model=model,
tokenizer=tokenizer,
truncation=True,
max_length=512,
device=torch.device("cuda" if torch.cuda.is_available() else "cpu"),
)
print(classifier("Sorry, but I can't assist with that."))