Downloads · 30 days
0
yeahrmek/arxiv-math-lean
arxiv-math-lean is a machine learning model from yeahrmek. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is a BPE tokenizer based on "Salesforce/codegen-350M-mono". The tokenizer has been trained to treat spaces like parts of the tokens (a bit like sentencepiece) so a word will be encoded differently whether it is a…
Downloads · 30 days
0
Access
Public
Updated Oct 27, 2022
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json2.9 MB · 86%
From the Hugging Face model README
This is a BPE tokenizer based on "Salesforce/codegen-350M-mono".
The tokenizer has been trained to treat spaces like parts of the tokens (a bit like sentencepiece)
so a word will be encoded differently whether it is at the beginning of the sentence (without space) or not.
We used ArXiv subset of The Pile dataset and proof steps from lean-step-public datasets to train the tokenizer.