Downloads · 30 days
13
6% of all-time downloads
omarmomen/transformer_base_final_2
transformer_base_final_2 is a fill-mask model from omarmomen. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
This model is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Struc…
Downloads · 30 days
13
6% of all-time downloads
All-time downloads
222
Public
Repo size
23.9 GB
Likes
0
Public
Click a slice to open those files.
.bin17.7 GB · 54%
From the Hugging Face model README
This model is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Structure Building" (https://aclanthology.org/2023.conll-babylm.29/)
<strong>omarmomen/transformer_base_final_2</strong> is a baseline vanilla transformer encoder.
The model is pretrained on the BabyLM 10M dataset using a custom pretrained RobertaTokenizer (https://huggingface.co/omarmomen/babylm_tokenizer_32k).