Downloads · 30 days
33
31% of all-time downloads
Nelera/ru-toxicity-detection
ru-toxicity-detection is a text classification model from Nelera. Use it when you need a label for a piece of text. It is set up for transformers.
Модель предназначена для классификации токсичности текста на русском языке. Обучена на основе архитектуры ai-forever/ru-en-RoSBERTa с использованием PyTorch Lightning.
Downloads · 30 days
33
31% of all-time downloads
All-time downloads
105
Public
Parameters
405M
1.6 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors1.6 GB · 99%
From the Hugging Face model README
Модель предназначена для классификации токсичности текста на русском языке.
Обучена на основе архитектуры ai-forever/ru-en-RoSBERTa с использованием PyTorch Lightning.
Модель решает задачу бинарной классификации текста:
0 – нейтральный, безопасный текст.1 – токсичный текст, содержащий грубость, мат, явные или скрытые оскорбления.from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("Nelera/ru-en-toxicity-rosberta")
model = AutoModelForSequenceClassification.from_pretrained("Nelera/ru-en-toxicity-rosberta")
text = "Ваш текст для проверки"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True, max_length=128)
with torch.no_grad():
logits = model(**inputs).logits
prob = torch.softmax(logits, dim=-1)[0][1].item()
pred = torch.argmax(logits, dim=1).item()
print(f"Класс: {pred} (Вероятность токсичности: {prob:.4f})")