Downloads · 30 days
4
5% of all-time downloads
TypicaAI/DarijaToxicityDetector
DarijaToxicityDetector is a text classification model from TypicaAI. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
DarijaToxicityDetector is a BERT-based binary text classifier that detects toxic / offensive content in Moroccan Darija (Moroccan Arabic dialect, written in Arabic script). It is fine-tuned from SI2M-Lab/DarijaBERT on…
Downloads · 30 days
4
5% of all-time downloads
All-time downloads
80
Public
Parameters
147M
590 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors590 MB · 100%
From the Hugging Face model README
DarijaToxicityDetector is a BERT-based binary text classifier that detects toxic / offensive content in Moroccan Darija (Moroccan Arabic dialect, written in Arabic script). It is fine-tuned from SI2M-Lab/DarijaBERT on the OMCD_Typica.ai_Mix dataset, a curated blend of the public OMCD dataset and Typica.ai's proprietary culturally grounded annotations.
The model is released by Typica.ai as part of its applied research on culturally localized AI for underserved languages, and is open-sourced for educational and research purposes.
📄 Companion paper: A Comparative Benchmark of a Moroccan Darija Toxicity Detection Model (Typica.ai) and Major LLM-Based Moderation APIs (OpenAI, Mistral, Anthropic) — the benchmark shows that this culturally adapted model outperforms general-purpose LLM moderation APIs on Moroccan Darija toxicity detection.
| Developed by | Hicham Assoudi — Typica.ai |
| Model type | BERT-based sequence classification (binary) |
| Language | Moroccan Darija (ary), Arabic script |
| Base model | SI2M-Lab/DarijaBERT |
| License | CC BY-NC 4.0 (non-commercial — education & research) |
| Paper | arXiv:2505.04640 |
| Contact | assoudi@typica.ai |
| id | label | meaning |
|---|---|---|
| 0 | clean | Non-toxic content |
| 1 | offensive | Toxic content (insults, hate, obscenity, culturally embedded aggression) |
Direct intended uses:
Out-of-scope uses:
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="TypicaAI/DarijaToxicityDetector",
)
texts = [
"هاد الفيديو ما عجبنيش بزاف", # expected: clean
"مول هاد الفيديو باسل وما مربّيش", # expected: offensive
"هاد الإنفلونسر مكلّخ غير كيخربق" # expected: offensive
]
print(classifier(texts, truncation=True, max_length=512))
# [{'label': 'clean', 'score': ...}, {'label': 'offensive', 'score': ...}]
The model was trained on OMCD_Typica.ai_Mix (12,758 Moroccan Darija comments), built as follows:
| Split | Examples | Used for |
|---|---|---|
| Train | 9,568 | Fine-tuning |
| Validation | 2,552 | Best-checkpoint selection |
| Test | 638 | Final evaluation & paper benchmark |
Each example carries sentence, label (ClassLabel: clean/offensive), idx, and origin (source provenance) fields.
The test split is publicly available for reproducibility in the benchmark GitHub repository. The proprietary training annotations are not released.
SI2M-Lab/DarijaBERTmax_length=512, dynamic padding (DataCollatorWithPadding)Hyperparameters:
| Hyperparameter | Value |
|---|---|
| Learning rate | 2e-5 |
| Train batch size | 16 |
| Eval batch size | 8 |
| Epochs | 10 |
| Weight decay | 0.01 |
| Eval/save strategy | per epoch, best model restored at end |
| Metric | Score |
|---|---|
| Accuracy | 0.8307 |
| Weighted F1 | 0.8308 |
Original benchmark — from the companion paper (May 2025), on the OMCD_Typica.ai_Mix test split (n = 630):
| Model | Accuracy | Macro F1 | Toxic F1 | Not-Toxic F1 |
|---|---|---|---|---|
| Typica.ai (this line of models) | 0.830 | 0.830 | 0.834 | 0.827 |
| OpenAI (omni-moderation-latest) | 0.652 | 0.644 | 0.589 | 0.699 |
| Mistral (mistral-moderation-latest) | 0.649 | 0.641 | 0.588 | 0.694 |
| Anthropic Claude (claude-3-haiku-20240307) | 0.659 | 0.617 | 0.743 | 0.492 |
Updated re-run — July 2026, same gold test set (n = 630, balanced), same inputs to all APIs, using each provider's then-current moderation endpoint (weighted precision / recall / F1):
| Model | Precision | Recall | F1-score |
|---|---|---|---|
| Typica.ai (custom BERT-based model) | 0.832 | 0.830 | 0.830 |
| Anthropic Claude (claude-haiku-4-5-20251001) | 0.695 | 0.657 | 0.646 |
| OpenAI (omni-moderation-latest) | 0.692 | 0.630 | 0.607 |
| Mistral (mistral-moderation-latest) | 0.633 | 0.592 | 0.571 |
Fourteen months after the original benchmark, the performance gap persists even against newer commercial models: the culturally adapted classifier still leads by ~18+ F1 points. General-purpose APIs continue to miss culturally nuanced toxicity (indirect insults, sarcasm, cultural idioms), while the specialized model maintains the best balance between catching toxic content and avoiding false positives.
This model deals with offensive content by design. It should support — not replace — human moderation. Misclassification can silence legitimate speech (false positives) or expose users to harm (false negatives). Deployers should implement human-in-the-loop review, appeal mechanisms, and threshold calibration appropriate to their community norms.
If you use this model, please cite:
@article{assoudi2025comparative,
title = {A Comparative Benchmark of a Moroccan Darija Toxicity Detection Model (Typica.ai) and Major LLM-Based Moderation APIs (OpenAI, Mistral, Anthropic)},
author = {Assoudi, Hicham},
journal = {arXiv preprint arXiv:2505.04640},
year = {2025},
url = {https://arxiv.org/abs/2505.04640}
}
Hicham Assoudi — Founder & Applied AI Researcher, Typica.ai · PhD (AI/NLP) Typica.ai — Independent applied research initiative 📧 assoudi@typica.ai · Linkedin . 🌐 typica.ai · 🤗 TypicaAI on Hugging Face