Downloads · 30 days
0
HuggingFaceGECLM/mix_tok_v2
mix_tok_v2 is a machine learning model from HuggingFaceGECLM. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
V1 of an English/code tokenizer. Byte-level BPE, 64k vocab, split digits (the difference with v1). Equal mix between: On the NL side: - Books - C4 - v1 of our CC (helen quality classifier) - enwiki - Gutenberg - Reddit
Downloads · 30 days
0
Access
Public
Updated Apr 13, 2023
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.7 MB · 100%
From the Hugging Face model README
V1 of an English/code tokenizer. Byte-level BPE, 64k vocab, split digits (the difference with v1). Equal mix between: On the NL side:
On the code side:
For a total of 1/3 code data (although there is a lot of English in Stackexchange and GH).