Downloads · 30 days
0
nilq/baby-tokenizer
baby-tokenizer is a machine learning model from nilq. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Compact sentencepiece tokenizer for sample-efficient English language modeling, simply tokenizing natural language.
Downloads · 30 days
0
Access
Public
Updated Feb 22, 2024
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json843 KB · 100%
From the Hugging Face model README
Compact sentencepiece tokenizer for sample-efficient English language modeling, simply tokenizing natural language.
from transformers import AutoTokenizer
tokenizer_baby = AutoTokenizer.from_pretrained("nilq/baby-tokenizer")
from tokenizers import Tokenizer
tokenizer_baby = Tokenizer.from_pretrained("nilq/baby-tokenizer")
This tokeniser is derived from the BabyLM 100M dataset of mixed domain data, consisting of the following sources: