Downloads · 30 days
8
24% of all-time downloads
HyeoniLEE/book_finetuned_model
book_finetuned_model is a fill-mask model from HyeoniLEE. Use it when you need the model to fill a missing word. It is set up for transformers.
QA task 및 finetuning 방법 교육 자료를 위해 bert model을 book 데이터셋에 맞춰 finetuning 한 모델입니다.
Downloads · 30 days
8
24% of all-time downloads
All-time downloads
33
Public
Parameters
110M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
QA task 및 finetuning 방법 교육 자료를 위해 bert model을 book 데이터셋에 맞춰 finetuning 한 모델입니다.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
from transformers import Trainer, TrainingArguments, AutoTokenizer, AutoModelForMaskedLM
from datasets import load_dataset
# 데이터셋 로드
dataset = load_dataset("HyeoniLEE/books_dataset")
# 데이터셋을 훈련 세트와 검증 세트로 나누기
dataset = dataset["train"].train_test_split(test_size=0.1) # 10%를 검증 세트로 사용
# 토크나이저 및 모델 로드
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
model = AutoModelForMaskedLM.from_pretrained("bert-base-uncased")
)
def preprocess_function(examples):
texts = []
for i in range(len(examples["title"])):
text = (
(examples["title"][i] if examples["title"][i] is not None else "") + " " +
(examples["author"][i] if examples["author"][i] is not None else "") + " " +
(examples["table_of_contents"][i] if examples["table_of_contents"][i] is not None else "") + " " +
(examples["book_intro"][i] if examples["book_intro"][i] is not None else "") + " " +
(examples["publisher_review"][i] if examples["publisher_review"][i] is not None else "") + " " +
(examples["review"][i] if examples["review"][i] is not None else "")
)
texts.append(text)
tokenized_inputs = tokenizer(texts, truncation=True, padding="max_length", max_length=512)
tokenized_inputs["labels"] = tokenized_inputs["input_ids"].copy() # labels 추가
return tokenized_inputs
# 데이터셋 전처리
tokenized_dataset = dataset.map(preprocess_function, batched=True)
training_args = TrainingArguments(
output_dir="./results",
evaluation_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=4,
per_device_eval_batch_size=4,
num_train_epochs=3,
weight_decay=0.01,
)
dataset = dataset["train"].train_test_split(test_size=0.1)
10%를 검증 데이터셋으로 사용

교육자료로 파인튜닝한 모델이라, 상당히 성능이 떨어지는 경향이 있다.
파인튜닝 설계 지점을 정확하게 계획하고, 디테일하게 만들어야 GPT와 같은 좋은 성능을 기대할 수 있다고 본다.
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]