Downloads · 30 days
15
44% of all-time downloads
quyle1304/PhoBERT-EmotionClassifier
PhoBERT-EmotionClassifier is a machine learning model from quyle1304. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Downloads · 30 days
15
44% of all-time downloads
All-time downloads
34
Public
Parameters
135M
540 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors540 MB · 100%
From the Hugging Face model README
Mô hình phân loại cảm xúc tiếng Việt (Emotion Classification) được fine-tune từ vinai/phobert-base-v2 trên bộ dữ liệu UIT-VSMEC (Vietnamese Students’ Multilabel Emotion Corpus).
Mô hình này được tối ưu hóa để nhận diện 01 cảm xúc chủ đạo (Single-label) và có khả năng hiểu ngữ nghĩa của các Icon/Emoji phổ biến.
vinai/phobert-base-v2Other, Disgust, Enjoyment, Sadness, Fear, Surprise, Anger.wseg).Kết quả đánh giá trên tập test của UIT-VSMEC (Single-label metrics):
| Metric | Score |
|---|---|
| Accuracy | 0.6580 |
| F1-Macro | 0.6291 |
| F1-Weighted | 0.6572 |
(Nguồn: Cell trong notebook)
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from vncorenlp import VnCoreNLP
import torch
import numpy as np
# 1. Load Model & Tokenizer
model_name = "quyle1304/PhoBERT-EmotionClassifier" # Thay bằng đường dẫn model của bạn
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# 2. Setup VnCoreNLP (Bắt buộc để tách từ)
# Tải VnCoreNLP-1.1.1.jar và cam/ tại: https://github.com/vncorenlp/VnCoreNLP
rdr = VnCoreNLP("path/to/VnCoreNLP-1.1.1.jar", annotators="wseg")
def segment(text):
return " ".join([" ".join(sent) for sent in rdr.tokenize(text)])
# 3. Predict Function
def predict(text):
text_seg = segment(text)
inputs = tokenizer(text_seg, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
probs = torch.softmax(logits, dim=1).cpu().numpy()[0] # Dùng Softmax cho Single-label
labels = ["Other", "Disgust", "Enjoyment", "Sadness", "Fear", "Surprise", "Anger"]
pred_label = labels[np.argmax(probs)]
confidence = np.max(probs)
return pred_label, confidence
# 4. Run
examples = [
"Hôm nay trời đẹp quá, mình cảm thấy rất vui! 😂",
"Phim này chán ngắt, phí cả tiền vé 😡"
]
for text in examples:
label, conf = predict(text)
print(f"Text: {text} | Emotion: {label} ({conf:.2%})")