Downloads · 30 days
23
13% of all-time downloads
perplexity-correlations/fasttext-lambada-es-target
fasttext-lambada-es-target is a machine learning model from perplexity-correlations. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for fasttext. The card lists the license as mit.
This is the fastText pretraining data filter targeting the LAMBADA ES task, discussed in the Perplexity Correlations paper: Improving Pretraining Data Using Perplexity Correlations. This filter uses perplexity correla…
Downloads · 30 days
23
13% of all-time downloads
All-time downloads
171
Public
Repo size
3.9 GB
Likes
0
Public
Click a slice to open those files.
.bin3.9 GB · 100%
From the Hugging Face model README
This is the fastText pretraining data filter targeting the LAMBADA ES task, discussed in the Perplexity Correlations paper: Improving Pretraining Data Using Perplexity Correlations. This filter uses perplexity correlations to identify high-quality pretraining data without requiring any LLM training.
This model is a data filter, not a language model itself, and should be used to select high-quality data for training LLMs. The filter works by estimating optimal weights for pretraining data selection based on the correlation between LLM perplexity on the data and downstream benchmark performance. It then uses these weights to train a fastText classifier, which can be used to filter new text data and select the highest quality samples.
Code: https://github.com/TristanThrush/perplexity-correlations
@misc{thrush2024perplexitycorrelations,
title={Improving Pretraining Data Using Perplexity Correlations},
author={Tristan Thrush and Christopher Potts and Tatsunori Hashimoto},
year={2024},
eprint={2409.05816},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2409.05816},
}