Downloads · 30 days
0
omarmomen/babylm_tokenizer_32k
babylm_tokenizer_32k is a machine learning model from omarmomen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This tokenizer is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit S…
Downloads · 30 days
0
Access
Public
Updated Mar 26, 2024
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.8 MB · 87%
From the Hugging Face model README
This tokenizer is part of the experiments in the published paper at the BabyLM workshop in CoNLL 2023. The paper titled "Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Structure Building" (https://aclanthology.org/2023.conll-babylm.29/)
<strong>omarmomen/babylm_tokenizer_32k</strong> is a RobertaTokenizer that is pretrained on the BabyLM 10M dataset (cased) with 32K tokens.