Downloads · 30 days
27
9% of all-time downloads
omarmomen/roberta_base_32k_final
roberta_base_32k_final is a fill-mask model from omarmomen. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
This model is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Struc…
Downloads · 30 days
27
9% of all-time downloads
All-time downloads
307
Public
Repo size
885 MB
Likes
0
Public
Click a slice to open those files.
.bin443 MB · 100%
From the Hugging Face model README
This model is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Structure Building" (https://aclanthology.org/2023.conll-babylm.29/)
<strong>omarmomen/roberta_base_32k_final</strong> is a baseline RobertaModel.
The model is pretrained on the BabyLM 10M dataset using a custom pretrained RobertaTokenizer (https://huggingface.co/omarmomen/babylm_tokenizer_32k).