Downloads · 30 days
0
bsgcasa/toxic-talking-detector
toxic-talking-detector is a text classification model from bsgcasa. Use it when you need a label for a piece of text. It is set up for onnxruntime. The card lists the license as apache-2.0.
A fine-tuned DistilBERT model for multi-label toxicity classification, exported to ONNX (INT8 quantized) for lightweight deployment.
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
67.4 MB
Likes
0
Public
Click a slice to open those files.
.onnx67.4 MB · 99%
From the Hugging Face model README
A fine-tuned DistilBERT model for multi-label toxicity classification, exported to ONNX (INT8 quantized) for lightweight deployment.
This model is fine-tuned from distilbert-base-uncased on the Jigsaw Unintended Bias in Toxicity Classification dataset. It predicts continuous toxicity scores (0-1) across 7 categories for a given text input.
Labels:
toxicitysevere_toxicityobscenethreatinsultidentity_attacksexual_explicitDesigned for conversation moderation systems — score individual messages or aggregate scores across a multi-turn conversation to detect escalating toxicity.
| File | Description |
|---|---|
model_int8.onnx | Quantized ONNX model (~67MB) |
vocab.txt | Tokenizer vocabulary |
tokenizer.json | Fast tokenizer config |
tokenizer_config.json | Tokenizer settings |
special_tokens_map.json | Special token mapping |
label_order.json | Output label order (maps model output indices to category names) |
distilbert-base-uncasedfrom huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
import onnxruntime as ort
import numpy as np
import json
REPO_ID = "bsgcasa/toxic-talking-detector"
model_path = hf_hub_download(repo_id=REPO_ID, filename="model_int8.onnx")
label_order_path = hf_hub_download(repo_id=REPO_ID, filename="label_order.json")
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
with open(label_order_path) as f:
LABELS = json.load(f)
session = ort.InferenceSession(model_path)
def predict(text: str) -> dict:
enc = tokenizer(text, return_tensors="np", padding="max_length", truncation=True, max_length=128)
ort_inputs = {
"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64),
}
logits = session.run(["logits"], ort_inputs)[0]
probs = 1 / (1 + np.exp(-logits))[0]
return {label: float(p) for label, p in zip(LABELS, probs)}
print(predict("You're completely useless, shut up."))
Apache 2.0