Downloads · 30 days
36
17% of all-time downloads
aref-j/emotion-classifier-bert-fa-v1
emotion-classifier-bert-fa-v1 is a text classification model from aref-j. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This is a fine-tuned BERT model for classifying emotions in Persian text, specifically detecting 6 emotion categories: ANGRY, FEAR, HAPPY, HATE, SAD, SURPRISE. It was developed using a merged dataset of Persian emotio…
Downloads · 30 days
36
17% of all-time downloads
All-time downloads
218
Public
Parameters
163M
651 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors651 MB · 99%
From the Hugging Face model README
This is a fine-tuned BERT model for classifying emotions in Persian text, specifically detecting 6 emotion categories: ANGRY, FEAR, HAPPY, HATE, SAD, SURPRISE. It was developed using a merged dataset of Persian emotion corpora and is designed for applications like sentiment analysis on Persian tweets.
This model is a fine-tuned version of ParsBERT (HooshvareLab/bert-base-parsbert-uncased) for emotion classification in Persian text. It uses a BERT base architecture with a sequence classification head to predict one of six emotion labels from input text. The model addresses class imbalance through weighted cross-entropy loss and was trained on a combined dataset of Persian tweets and short texts.
Use the code below to get started with the model.
from transformers import AutoTokenizer, AutoModelForSequenceClassification, pipeline
model_name = "aref-j/emotion-classifier-bert-fa-v1"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Create the classification pipeline
classifier = pipeline("text-classification", model=model, tokenizer=tokenizer)
# Example usage
result = classifier("چه هوای زیبایی امروز است")
print(result) # e.g. [{'label': 'HAPPY', 'score': 0.99}]
The model was trained on a merged dataset from three Persian emotion corpora:
Datasets were standardized, cleaned (normalization with Parsivar, removal of URLs, mentions, emojis, etc.), deduplicated, and split into 90% train / 10% validation, with ArmanEmo held out for testing.
Text was normalized using Parsivar, with character mapping, diacritic removal, and stripping of URLs, mentions, hashtags, emojis, punctuation, digits, and extra spaces. Multi-label instances in EmoPars were converted to single-label via dominant label.
Held-out ArmanEmo test set.
Evaluation disaggregated by emotion classes (ANGRY, FEAR, HAPPY, HATE, SAD, SURPRISE).
Accuracy (overall correct predictions), Macro F1-score (average F1 across classes, treating all equally), Precision, Recall, and Confusion Matrix.
Detailed per-class metrics and confusion matrix available in the repository.
BibTeX:
@misc{jafary2023persianemotion,
author = {Aref Jafary},
title = {Persian Emotion Classification with BERT},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/ArefJafary/Persian-Emotion-Classification-BERT}}
}
APA: Jafary, A. (2023). Persian Emotion Classification with BERT [Repository]. GitHub. https://github.com/ArefJafary/Persian-Emotion-Classification-BERT
Contact via GitHub: https://github.com/ArefJafary