Downloads · 30 days
5
1% of all-time downloads
kms7530/roberta-base-infringement-detect
roberta-base-infringement-detect is a text classification model from kms7530. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
klue/roberta-base 모델을 이용하여, 두 컨텐츠간의 유사여부를 확인하는 모델입니다.
Downloads · 30 days
5
1% of all-time downloads
All-time downloads
410
Public
Parameters
111M
1.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.onnx555 MB · 56%
From the Hugging Face model README
klue/roberta-base 모델을 이용하여, 두 컨텐츠간의 유사여부를 확인하는 모델입니다.
자체구축된 1,310개의 참인 유사 컨텐츠 쌍을 이용하여, 셔플 후 참/거짓 비율 1:2인 데이터셋을 생성하여 학습시켰습니다.
이외의 학습시 파라미터는 다음과 같습니다.
| Parameter | Value |
|---|---|
train_batch_size | 16 |
num_train_epochs | 5 |
weight_decay | 0.01 |
learning_rate | 2e-5 |
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_name = "kms7530/roberta-base-infringement-detect"
model = AutoModelForSequenceClassification.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)
모델에 추론 시 다음과 같이 입력해야 합니다.
[CLS]\
[unused0]<ORIGINAL_CONTENT_TITLE>\
[unused1]<ORIGINAL_CONTENT>[SEP] \
[unused0]<TEST_CONTENT_TITLE>\
[unused1]<TEST_CONTENT>[SEP]