Downloads · 30 days
17
5% of all-time downloads
Builder117/distilbert-insecure-output
distilbert-insecure-output is a text classification model from Builder117. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
Fine-tuned DistilBERT classifier that detects dangerous payloads in LLM-generated output.
Downloads · 30 days
17
5% of all-time downloads
All-time downloads
372
Public
Parameters
67M
4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
Fine-tuned DistilBERT classifier that detects dangerous payloads in LLM-generated output.
Covers OWASP LLM Top 10 — LLM02: Insecure Output Handling.
Malicious code or injection payloads that an LLM might generate, including:
<script>alert(document.cookie)</script>'; DROP TABLE users; --| cat /etc/passwd../../etc/shadow| Label | ID | Meaning |
|---|---|---|
SAFE | 0 | Safe output (normal text, parameterized queries, sanitized code) |
MALICIOUS | 1 | Dangerous payload detected |
from transformers import pipeline
clf = pipeline("text-classification", model="Builder117/distilbert-insecure-output")
clf("<script>alert(document.cookie)</script>")
# [{'label': 'MALICIOUS', 'score': 0.98}]
clf("SELECT * FROM products WHERE id = ?")
# [{'label': 'SAFE', 'score': 0.97}] # parameterized — safe
distilbert-base-uncasedLLM Threat Shield — OWASP LLM Top 10 detection suite.