Downloads · 30 days
18
13% of all-time downloads
0khacha/darija-toxicity-classifier
darija-toxicity-classifier is a text classification model from 0khacha. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
A transformer-based NLP model for detecting toxic content in Moroccan Darija and Arabizi.
Downloads · 30 days
18
13% of all-time downloads
All-time downloads
136
Public
Parameters
171M
2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors682 MB · 100%
From the Hugging Face model README
A transformer-based NLP model for detecting toxic content in Moroccan Darija and Arabizi.
This model is specifically designed to handle the linguistic complexity of Moroccan dialect, including Arabizi (Arabic written in Latin characters with numbers) such as:
3 → ع7 → ح9 → قIt also supports code-switched text mixing Darija, Arabic, French, English, and Tamazight.
| Property | Value |
|---|---|
| Model ID | 0khacha/darija-toxicity-classifier |
| Architecture | Fine-tuned from SI2M-Lab/DarijaBERT-arabizi |
| Task | Binary Sequence Classification (Safe / Toxic) |
| Framework | Hugging Face Transformers |
| Training Data | 16,000+ labeled Moroccan Darija/Arabizi samples |
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="0khacha/darija-toxicity-classifier"
)
result = classifier("salam khouya")
print(result)
# Output: [{'label': 'SAFE', 'score': 0.9845}]
Built specifically for Moroccan linguistic patterns — not generic Arabic.
Understands numeric character substitutions like:
in3alsa7a3likomThe model was trained with specialized normalization:
w-a-l-o → walo)n 3 a l → n3al)heeeey → hey)| Metric | Score |
|---|---|
| Accuracy | ~94% |
| F1-Score | ~93% |
| Inference Speed (GPU) | ~50ms |
Note: Performance may vary depending on hardware and deployment setup.
Input:
"bghit nakol"
Output:
Safe (98.45%)
Input:
"rak stupid"
Output:
Toxic
The training dataset is not publicly available for privacy and ethical reasons.
For research collaboration: 📩 [email protected]
MIT License
If you use this model in your research, please cite:
@misc{darija-toxicity-classifier,
author = {Khacha, Mohamed},
title = {Darija Toxicity Classifier},
year = {2024},
publisher = {HuggingFace},
url = {https://huggingface.co/0khacha/darija-toxicity-classifier}
}
Contributions, issues, and feature requests are welcome!
Feel free to check the issues page.
Made with ❤️ for the Moroccan NLP community