Downloads · 30 days
14
10% of all-time downloads
perplexity-correlations/fasttext-lambada-it-target
fasttext-lambada-it-target is a machine learning model from perplexity-correlations. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for fasttext. The card lists the license as mit.
This fastText model is a filter for selecting high-quality pretraining data, as described in Improving Pretraining Data Using Perplexity Correlations. It targets the LAMBADA IT task.
Downloads · 30 days
14
10% of all-time downloads
All-time downloads
136
Public
Repo size
3.9 GB
Likes
0
Public
Click a slice to open those files.
.bin3.9 GB · 100%
From the Hugging Face model README
This fastText model is a filter for selecting high-quality pretraining data, as described in Improving Pretraining Data Using Perplexity Correlations. It targets the LAMBADA IT task.
The model uses perplexity correlations to identify text segments highly correlated with strong performance on downstream benchmarks. It doesn't perform text classification directly; instead, it outputs a score indicating the suitability of a text segment for pretraining.
For complete usage instructions and the theoretical background, please refer to the project's GitHub repository.