Downloads · 30 days
15
2% of all-time downloads
hazyresearch/mamba-1b
mamba-1b is a machine learning model from hazyresearch. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This model is pretrained Mamba model. The goal of this model is to provide a quality reference for the Based architecture.
Downloads · 30 days
15
2% of all-time downloads
All-time downloads
727
Public
Repo size
10.6 GB
Likes
1
Public
Click a slice to open those files.
.bin5.3 GB · 100%
From the Hugging Face model README
This model is pretrained Mamba model. The goal of this model is to provide a quality reference for the Based architecture.
As a quality reference, we include a pretrained Attention (Llama architecture) model provided here: https://huggingface.co/hazyresearch/attn-1b, and a pretrained Based model provided here: https://huggingface.co/hazyresearch/based-1b
All three checkpoints are pretrained on 10Bn tokens of the Pile in the exact same data order using next token prediction.
The model implementation and training code that produced the model are provided here: https://github.com/HazyResearch/based
The purpose of this work is to evaluate the language modeling quality of a new efficient architecture, Based.
We include a series of benchmarks that you can use to evaluate quality:
Please consider citing this paper if you use our work:
@article{arora2024simple,
title={Simple linear attention language models balance the recall-throughput tradeoff},
author={Arora, Simran and Eyuboglu, Sabri and Zhang, Michael and Timalsina, Aman and Alberti, Silas and Zinsley, Dylan and Zou, James and Rudra, Atri and Ré, Christopher},
journal={arXiv:2402.18668},
year={2024}
}
Please reach out to [email protected], [email protected], and [email protected] with questions.