Downloads · 30 days
15
4% of all-time downloads
imaneumabderahmane/Arabertv02-classifier-FA
Arabertv02-classifier-FA is a text classification model from imaneumabderahmane. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
FA-AraBert is an Arabic binary text classification model designed to detect whether a user query is related to first-aid. The classifier serves as the intent detection and safety filtering component of an MSA first-ai…
Downloads · 30 days
15
4% of all-time downloads
All-time downloads
410
Public
Parameters
135M
541 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors541 MB · 99%
From the Hugging Face model README
FA-AraBert is an Arabic binary text classification model designed to detect whether a user query is related to first-aid. The classifier serves as the intent detection and safety filtering component of an MSA first-aid chatbot pipeline. Two models are developed and evaluated in this project: FA-AraBERTv2 and FA-AraBERTv0.2, both fine-tuned on the FALAH-Mix dataset (1,028 question–answer (QA) pairs, including 924 non first-aid pairs and 104 first-aid pairs) based on AraBERT base models. These classifiers were systematically compared under multiple training configurations to identify the most suitable model for deployment.
The FA-AraBERT classifier was developed as part of the PFE project titled:
Towards Building an Arabic First-Aid Chatbot using FA-AraBERT Classifier and FALAH Dataset.
Shared by: MABROUK Imane
Model type: Transformer-based text classification model (BERT architecture)
Language(s) (NLP): MSA(Modern Standard Arabic)
License: Apache 2.0
Finetuned from model: AraBERTv02
This model can be used directly for:
The model can be integrated into:
This model is not suitable for:
Use the code below to get started with the model:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "imaneumabderahmane/Arabertv02-classifier-FA"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "ما هي الإسعافات الأولية لحروق الدرجة الأولى؟"
inputs = tokenizer(text, return_tensors="pt", truncation=True)
with torch.no_grad():
outputs = model(**inputs)
prediction = torch.argmax(outputs.logits, dim=-1).item()
print(prediction)
The FA-AraBERTv02 classifier was trained and evaluated on the FALAH-Mix dataset, which contains 1,028 Arabic question–answer pairs (924 non–first-aid QA pairs and 104 first-aid QA pairs). The dataset exhibits a strong class imbalance, with approximately 90% non–first-aid queries and 10% first-aid queries. The data was split into training, development, and test sets while preserving the original class distribution.
For more details about the FALAH and FALAH-Mix datasets, developed as part of the PFE project, please refer to: https://huggingface.co/datasets/imaneumabderahmane/FALAH
To mitigate class imbalance, the training set was augmented with additional first-aid samples from external datasets, including the Mayo Clinic First-Aid dataset (374 first-aid QA pairs) and the 68 first-aid QA pairs from the AHD dataset. This resulted in a balanced training set of 1,184 samples, while the development and test sets remained unchanged.
The following table presents the FALAH-Mix dataset before balancing the training set:
<p align="center"> <img src="./distribution of the FALAH-Mix dataset.png" width="900"/> </p>The following table presents the FALAH-Mix dataset after balancing the training set:
<p align="center"> <img src="./table FALAH-Mix training set balanced.png" width="700"/> </p>Note: The FALAH-Mix dataset was split according to the emergency labels. As a result, the training, development, and test sets each contain approximately 90% non–first-aid QA pairs and 10% first-aid QA pairs. The balancing strategy was applied only to the training set in order to evaluate its impact on the classifier.
The models were fine-tuned using supervised learning with the following configuration:
The test split from the FALAH-Mix dataset: https://huggingface.co/datasets/imaneumabderahmane/FALAH
The following Table summarizes the Macro F1 scores obtained by FA-AraBERTv2 and FA-AraBERTv0.2 under different training configurations.
<p align="center"> <img src="./table results classifier arabert.png" width="800"/> </p> The best performance was achieved when fine-tuning on the balanced FALAH-Mix training set with class weighting. FA-AraBERTv2 achieved a Macro F1-score of 0.6379, slightly outperforming FA-AraBERTv0.2. Due to this consistent advantage, FA-AraBERTv2 was selected for deployment in the final chatbot system. It should be noted that statistical significance testing was not performed due to time constraints related to the project.