Downloads · 30 days
33
4% of all-time downloads
b4c0n/KAi-Toxicity-Filter
KAi-Toxicity-Filter is a text classification model from b4c0n. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
日本語の有害表現検出に特化したモデル Japanese toxicity detection model specialized for Japanese language
Downloads · 30 days
33
4% of all-time downloads
All-time downloads
796
Public
Parameters
111M
1.3 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors445 MB · 100%
From the Hugging Face model README
日本語の有害表現検出に特化したモデル
Japanese toxicity detection model specialized for Japanese language
日本語テキストを有害/非有害に分類するモデルです。このモデルはtohoku-nlp/bert-base-japanese-v3をベースに、日本語の有害表現検出タスクでファインチューニングされています。
以下のデータで学習されています:
検証データセットでの評価結果:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "b4c0n/KAi-Toxicity-Filter"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "終わってる暴言"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
probs = torch.softmax(outputs.logits, dim=1)
toxic_prob = probs[0][1].item()
print(f"有害確率: {toxic_prob:.2%}")
KAi (かい鯖グループAI) における日本語テキストの有害コンテンツ検出・フィルタリングのために開発されました。
主な用途:
⚠️ このモデルは有害コンテンツデータで学習されています。責任を持って使用してください。
Apache 2.0
このモデルは inspection-ai/japanese-toxic-dataset (Apache 2.0 License) のデータを使用しています。
This model classifies Japanese text as toxic or non-toxic. It is fine-tuned from tohoku-nlp/bert-base-japanese-v3 for Japanese toxicity detection tasks.
This model was trained on:
Evaluation results on validation dataset:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "b4c0n/KAi-Toxicity-Filter"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "toxic expression"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
probs = torch.softmax(outputs.logits, dim=1)
toxic_prob = probs[0][1].item()
print(f"Toxic probability: {toxic_prob:.2%}")
This model was developed for the KAi (KaisabaGroupAI) to detect and filter harmful content in Japanese text.
Primary Use Cases:
⚠️ This model was trained on toxic content data. Please use responsibly.
Apache 2.0
@misc{kai-toxicity-filter,
author = {b4c0n},
title = {KAi Toxicity Filter: Japanese Toxicity Detection Model},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/b4c0n/KAi-Toxicity-Filter}}
}
This model uses data from inspection-ai/japanese-toxic-dataset (Apache 2.0 License).