Downloads · 30 days
33
22% of all-time downloads
thejosango/nuha-ajp-binary
nuha-ajp-binary is a text classification model from thejosango. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
nuha-ajp-binary is a binary Arabic text classifier that detects hate speech in Jordanian social media comments. It fine-tunes nuha-ajp-mlm — a domain-adapted Arabic BERT — and outputs one of two labels:
Downloads · 30 days
33
22% of all-time downloads
All-time downloads
149
Public
Parameters
78.5M
22 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin314 MB · 50%
From the Hugging Face model README
nuha-ajp-binary is a binary Arabic text classifier that detects hate speech in Jordanian social media comments. It fine-tunes nuha-ajp-mlm — a domain-adapted Arabic BERT — and outputs one of two labels:
| Label | Meaning |
|---|---|
non-hate-speech | Not Online Violence |
hate-speech | Offensive Language or Online Gender Based Violence |
This model was developed as part of a pilot proof-of-concept for the NUHA project by the Jordan Open Source Association (JOSA). Performance metrics reflect the complexity of hate speech detection in colloquial Arabic and the exploratory nature of this initial effort.
For a more granular three-class classifier, see nuha-ajp-trinary.
Classifying Arabic social media comments as hate speech or non-hate speech, particularly for Jordanian Arabic content from Facebook and X (Twitter).
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="thejosango/nuha-ajp-binary",
tokenizer="thejosango/nuha-ajp-binary",
)
result = classifier("أنتِ امرأة رائعة")
print(result)
# [{'label': 'non-hate-speech', 'score': ...}]
For batch inference:
comments = ["أنتِ امرأة رائعة", "اخرسي يا غبية"]
results = classifier(comments)
for comment, result in zip(comments, results):
print(f"{result['label']} ({result['score']:.2f}): {comment}")
Fine-tuned on the binary configuration of thejosango/nuha-ajp-dataset, which maps:
non-hate-speechhate-speechhate-speechAt training and inference time, the following normalisation is applied to input text (in addition to the dataset-level Arabic-only filtering):
[رابط] token[مستخدم] token[بريد] token| Parameter | Value |
|---|---|
| Base model | thejosango/nuha-ajp-mlm |
| Hidden layers | 4 (reduced from base's 12) |
| Classifier dropout | 0.50 |
| Learning rate | 5e-5 |
| LR schedule | Linear |
| Batch size | 64 |
| Epochs | 5 |
| Weight decay | 1e-3 |
| Label smoothing | 0.1 |
| Weighted loss | Yes (balanced class weights) |
| Data augmentation | Yes (contextual word substitution, ratio 0.75) |
| Framework | Transformers 4.32.1, PyTorch 2.0.1 |
Evaluated on the validation split of thejosango/nuha-ajp-dataset (binary configuration):
| Metric | Value |
|---|---|
| F1 | 0.6879 |
| Precision | 0.6464 |
| Recall | 0.7351 |
| Loss | 0.5743 |
This model was developed as part of an initial pilot study. Results should be interpreted accordingly.