Downloads · 30 days
37
16% of all-time downloads
BAR-ILAN/hebrew-toxicity-detector
hebrew-toxicity-detector is a machine learning model from BAR-ILAN. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A fine-tuned BERT-based model for detecting toxic and offensive text, with a focus on Hebrew social media comments.
Downloads · 30 days
37
16% of all-time downloads
All-time downloads
226
Public
Parameters
178M
711 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors711 MB · 100%
From the Hugging Face model README
A fine-tuned BERT-based model for detecting toxic and offensive text, with a focus on Hebrew social media comments.
This model was developed as part of my final Computer Science project.
The goal is to improve automatic detection of harmful, offensive, and toxic comments, especially in Hebrew, where existing toxicity models often perform less accurately.
The model is based on a Transformer text-classification architecture and was fine-tuned on a custom dataset of toxic and non-toxic examples.
It is designed to classify short user-generated text such as:
This model is used as part of a browser extension that helps reduce direct exposure to harmful comments by blurring content detected as toxic.
model.safetensors — fine-tuned model weightsconfig.json — model configurationtokenizer.json — tokenizer filetokenizer_config.json — tokenizer configurationfrom transformers import pipeline
model_path = "maayan3330/hebrew-toxicity-detector"
classifier = pipeline(
"text-classification",
model=model_path,
tokenizer=model_path
)
result = classifier("אתה פשוט מגעיל")
print(result)