Downloads · 30 days
51
1% of all-time downloads
rkdaldus/ko-sent5-classification
ko-sent5-classification is a feature extraction model from rkdaldus. Use it when you need embeddings to search or compare text. It is set up for transformers.
이 프로젝트는 한국어 텍스트의 감정을 분류하는 KoBERT 기반의 감정 분류 모델을 학습하고 활용하는 코드를 포함합니다. 이 모델은 입력된 텍스트가 분노(Anger), 두려움(Fear), 기쁨(Happy), 평온(Tender), 슬픔(Sad) 중 어떤 감정에 해당하는지를 예측합니다.
Downloads · 30 days
51
1% of all-time downloads
All-time downloads
5.4K
Public
Parameters
92.2M
369 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors369 MB · 100%
From the Hugging Face model README
이 프로젝트는 한국어 텍스트의 감정을 분류하는 KoBERT 기반의 감정 분류 모델을 학습하고 활용하는 코드를 포함합니다. 이 모델은 입력된 텍스트가 분노(Anger), 두려움(Fear), 기쁨(Happy), 평온(Tender), 슬픔(Sad) 중 어떤 감정에 해당하는지를 예측합니다.
필요 라이브러리 설치:
transformers, datasets, torch, pandas, scikit-learn 라이브러리를 설치합니다.
데이터 불러오기: ai hub 에 등록된 한국어 감성 대화 데이터로부터 감정 분류용 CSV 파일을 불러옵니다.
데이터셋 준비:
Dataset으로 변환.label_int 컬럼을 labels로 변경.monologg/kobert 토크나이저를 이용해 입력 텍스트를 토큰화.input_ids, attention_mask, labels만 남겨 학습 준비 완료.모델 및 학습 설정:
monologg/kobert 모델을 불러와 5개의 감정 레이블을 분류하도록 설정.learning_rate=2e-5, num_train_epochs=10, batch_size=16.학습 진행 및 모델 저장:
monologg/kobert 기반이며, 분류 레이블은 다음과 같습니다:
단순 문장 입력 감정 분석:
엑셀 파일에서 감정 분석:
# 토크나이저 및 모델 로드
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# KoBERT 토크나이저와 모델 로드
tokenizer = AutoTokenizer.from_pretrained("monologg/kobert", trust_remote_code=True)
model = AutoModelForSequenceClassification.from_pretrained("rkdaldus/ko-sent5-classification")
# 사용자 입력 텍스트 감정 분석
text = "오늘 정말 행복해!"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
outputs = model(**inputs)
predicted_label = torch.argmax(outputs.logits, dim=1).item()
# 감정 레이블 정의
emotion_labels = {
0: ("Angry", "😡"),
1: ("Fear", "😨"),
2: ("Happy", "😊"),
3: ("Tender", "🥰"),
4: ("Sad", "😢")
}
# 예측된 감정 출력
print(f"예측된 감정: {emotion_labels[predicted_label][0]} {emotion_labels[predicted_label][1]}")