Downloads · 30 days
27
1% of all-time downloads
hazyresearch/based-1b-50b
based-1b-50b is a machine learning model from hazyresearch. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Downloads · 30 days
27
1% of all-time downloads
All-time downloads
2.5K
Public
Repo size
10.8 GB
Likes
1
Public
Click a slice to open those files.
.bin5.4 GB · 100%
From the Hugging Face model README
This model is pretrained Based model.
As a quality reference, we include a pretrained Mamba model provided here: https://huggingface.co/hazyresearch/mamba-1b-50b and a pretrained attention (Llama architecture) model provided here: https://huggingface.co/hazyresearch/attn-1b-50bn
All three checkpoints are pretrained on 50Bn tokens of the Pile in the exact same data order using next token prediction.
A WandB report for training is here: https://api.wandb.ai/links/hazy-research/ggo9rst2
The model implementation and training code that produced the model are provided here: https://github.com/HazyResearch/based
The purpose of this work is to evaluate the language modeling quality of a new efficient architecture, Based.
We include a series of benchmarks that you can use to evaluate quality:
Please consider citing this paper if you use our work:
@article{arora2024simple,
title={Simple linear attention language models balance the recall-throughput tradeoff},
author={Arora, Simran and Eyuboglu, Sabri and Zhang, Michael and Timalsina, Aman and Alberti, Silas and Zinsley, Dylan and Zou, James and Rudra, Atri and Ré, Christopher},
journal={arXiv:2402.18668},
year={2024}
}
Please reach out to [email protected], [email protected], and [email protected] with questions.