Downloads · 30 days
8
1% of all-time downloads
perplexity-correlations/fasttext-arc-easy-target
fasttext-arc-easy-target is a text classification model from perplexity-correlations. Use it when you need a label for a piece of text. It is set up for fasttext. The card lists the license as mit.
This fastText model is a pretraining data filter, targeting the ARC Easy task. It's designed to select high-quality pretraining data using perplexity correlations, as described in Improving Pretraining Data Using Perp…
Downloads · 30 days
8
1% of all-time downloads
All-time downloads
776
Public
Repo size
3.9 GB
Likes
0
Public
Click a slice to open those files.
.bin3.9 GB · 100%
From the Hugging Face model README
This fastText model is a pretraining data filter, targeting the ARC Easy task. It's designed to select high-quality pretraining data using perplexity correlations, as described in Improving Pretraining Data Using Perplexity Correlations. The model classifies text as either "include" or "exclude" for use in pretraining a language model. It does not itself represent a pretrained language model.
The filter was created using a method that leverages correlations between LLM losses on various texts and downstream benchmark performance. By selecting texts with high correlation, this model aims to improve the efficiency of the data selection process for pretraining LLMs.
Code: https://github.com/TristanThrush/perplexity-correlations