Downloads · 30 days
212
1% of all-time downloads
vaiv/kobigbird-roberta-large
kobigbird-roberta-large is a fill-mask model from vaiv. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
This is a large-sized Korean BigBird model introduced in our paper. The model draws heavily from the parameters of klue/roberta-large to ensure high performance. By employing the BigBird architecture and incorporating…
Downloads · 30 days
212
1% of all-time downloads
All-time downloads
22.2K
Public
Parameters
341M
2.7 GB on disk
Likes
6
Public
Click a slice to open those files.
.bin1.4 GB · 50%
How the weights are stored.
F32341M · 100%
From the Hugging Face model README
This is a large-sized Korean BigBird model introduced in our paper. The model draws heavily from the parameters of klue/roberta-large to ensure high performance. By employing the BigBird architecture and incorporating the newly proposed TAPER, the language model accommodates even longer input lengths.
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("vaiv/kobigbird-roberta-large")
model = AutoModelForMaskedLM.from_pretrained("vaiv/kobigbird-roberta-large")

Measurement on validation sets of the KLUE benchmark datasets

While our model achieves great results even without additional pretraining, further pretraining can refine the positional representations more.
@article{yang2023kobigbird,
title={KoBigBird-large: Transformation of Transformer for Korean Language Understanding},
author={Yang, Kisu and Jang, Yoonna and Lee, Taewoo and Seong, Jinwoo and Lee, Hyungjin and Jang, Hwanseok and Lim, Heuiseok},
journal={arXiv preprint arXiv:2309.10339},
year={2023}
}