Downloads · 30 days
15
10% of all-time downloads
WebOrganizer/LM-1b_1x-Sampling_over_Topics_for_MMLU
LM-1b_1x-Sampling_over_Topics_for_MMLU is a text generation model from WebOrganizer. Use it when you need the model to write or continue text. It is set up for transformers.
Downloads · 30 days
15
10% of all-time downloads
All-time downloads
145
Public
Parameters
1.4B
23.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt17.3 GB · 73%
From the Hugging Face model README
A 1.4B parameter model trained for 29B tokens from WebOrganizer/Corpus-200B.
The training data for this model was selected via:
Besides the HuggingFace model and tokenizer, the repository contains:
open_lm/: Contains the OpenLM config and final checkpointevals/: Evaluation results for various benchmarks
core_9mcqa/: Results of 9 multiple choice QA tasks with the OLMES evaluation frameworkmmlu/: MMLU results with the OLMES evaluation frameworkdclm/: Results using the DCLM evaluation frameworkperplexity/: Perplexity results using the huggingface trainerindices.tar.zst: The indices for the selected documents in each shard of the Corpus-200B dataset used for training. The indices can be extracted with tar --use-compress-program "zstd" -xf indices.tar.zst.To use this model, you need to install the open_lm library and add from open_lm.hf import * before loading the model with AutoModel.from_pretrained(...).
@article{wettig2025organize,
title={Organize the Web: Constructing Domains Enhances Pre-Training Data Curation},
author={Alexander Wettig and Kyle Lo and Sewon Min and Hannaneh Hajishirzi and Danqi Chen and Luca Soldaini},
journal={arXiv preprint arXiv:2502.10341},
year={2025}
}