Downloads · 30 days
193
23% of all-time downloads
youssefreda9/HAYAA
HAYAA is a text classification model from youssefreda9. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
[](https://huggingface.co/datasets/youssefreda9/HAYAA) [](https://github.com/youssefreda10/HAYAA)
Downloads · 30 days
193
23% of all-time downloads
All-time downloads
830
Public
Parameters
163M
1.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors651 MB · 100%
From the Hugging Face model README
Hayā is a fine-tuned UBC-NLP/MARBERTv2 model for binary Arabic toxicity classification (Safe / Toxic). It is designed to detect offensive language, hate speech, cyberbullying, profanity, and other forms of toxic content across all major Arabic dialects.
Hayā is the core classifier in a larger defense-in-depth content moderation system that includes rule-based layers for explicit profanity and a Chrome extension that blurs toxic content in real time.
Evaluated on a held-out test set of 99,759 sentences (stratified, zero data leakage):
| Metric | Score |
|---|---|
| Accuracy | 97.84% |
| F1 (Toxic) | 94.12% |
| F1 (Safe) | 98.45% |
Note on the numbers: Manual error analysis showed the model frequently outperformed the original human annotations — many counted "errors" were actually mislabels in the source datasets. Real-world performance on correctly-labeled data is therefore higher than the raw scores suggest.
from transformers import pipeline
classifier = pipeline("text-classification", model="youssefreda9/HAYAA", top_k=None)
results = classifier("أنت إنسان رائع ومحترم")
print(results)
# [[{'label': 'Safe', 'score': 0.99}, {'label': 'Toxic', 'score': 0.01}]]
results = classifier("يا حمار أنت")
print(results)
# [[{'label': 'Toxic', 'score': 0.98}, {'label': 'Safe', 'score': 0.02}]]
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("youssefreda9/HAYAA")
model = AutoModelForSequenceClassification.from_pretrained("youssefreda9/HAYAA")
text = "كلامك جميل جدا"
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True, padding=True)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.softmax(outputs.logits, dim=-1)
pred = torch.argmax(probs, dim=-1).item()
labels = {0: "Safe", 1: "Toxic"}
print(f"Prediction: {labels[pred]} ({probs[0][pred]:.2%})")
from transformers import pipeline
classifier = pipeline("text-classification", model="youssefreda9/HAYAA")
texts = [
"صباح الخير يا أصدقاء",
"أنت واطي ومحترمش حد",
"الجو جميل النهارده",
]
results = classifier(texts)
for text, result in zip(texts, results):
print(f"{result['label']} ({result['score']:.2%}): {text}")
| Parameter | Value |
|---|---|
| Max sequence length | 128 |
| Training Strategy | 2-Stage (4 epochs base corpus + 3 epochs hard edge-cases) |
| Batch size | 16 (effective: 32 with gradient accumulation) |
| Learning rate | 2e-5 (Stage 1), 5e-6 (Stage 2) |
| Warmup ratio | 0.1 |
| Weight decay | 0.01 |
| FP16 | ✅ |
| Loss | Weighted cross-entropy (class imbalance) |
| Best model metric | F1 (weighted) |
| Optimizer | AdamW |
| Seed | 42 |
The training set is imbalanced (more Safe than Toxic). A weighted cross-entropy loss was used, with the Toxic class weight dynamically computed as safe_count / toxic_count to ensure the model doesn't under-predict toxicity.
The Toxic label covers:
| Category | Examples |
|---|---|
| Profanity | Explicit swear words across all dialects |
| Hate speech | Incitement against groups based on religion, ethnicity, nationality |
| Cyberbullying | Personal insults, harassment, directed attacks |
| Racism | Racial slurs and discriminatory language |
| Sexism | Misogynistic and sexually explicit language |
| Religious hate | Blasphemy, sectarian attacks |
| Obfuscated toxicity | Intentional typos, spaced letters, homoglyphs |
| Dialect | Coverage |
|---|---|
| Egyptian | ✅ |
| Levantine (Syrian, Lebanese, Jordanian, Palestinian) | ✅ |
| Gulf (Saudi, Emirati, Kuwaiti, Bahraini, Omani, Qatari) | ✅ |
| Maghrebi (Moroccan, Algerian, Tunisian, Libyan) | ✅ |
| Iraqi | ✅ |
| Sudanese | ✅ |
| Modern Standard Arabic (MSA) | ✅ |
This model is one layer in the Hayā defense-in-depth pipeline:
| Layer | Function |
|---|---|
| L0 — Sanitize | Unicode normalization, homoglyph folding, emoji analysis |
| L1 — Dictionary | Context-aware instant matching (100% precision) |
| L1.5 — De-obfuscation | Resolves masked/spaced evasion attempts |
| L2 — This Model | Deep-learning classification for implicit toxicity |
The full pipeline ships as a Chrome Extension (Manifest V3) with PIN-protected parental controls. See the GitHub repository for the complete system.
@misc{hayaa2026,
title={Hayā: Fine-tuned MARBERTv2 for Multi-Dialect Arabic Toxicity Classification},
author={Youssef Reda},
year={2026},
url={https://huggingface.co/youssefreda9/HAYAA},
}
This model is released under the Apache 2.0 License.