Downloads · 30 days
119
6% of all-time downloads
Zyphra/Mamba-370M
Mamba-370M is a machine learning model from Zyphra. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Install Zyphra's fork of mamba-ssm (https://github.com/Zyphra/mamba) 1. git clone https://github.com/Zyphra/mamba.git 2. cd mamba 3. You need to install from source: pip install . or pip install -e . for editable mode…
Downloads · 30 days
119
6% of all-time downloads
All-time downloads
2K
Public
Repo size
60.1 GB
Likes
8
Public
Click a slice to open those files.
.bin55 GB · 100%
From the Hugging Face model README
Install Zyphra's fork of mamba-ssm (https://github.com/Zyphra/mamba)
git clone https://github.com/Zyphra/mamba.gitcd mambapip install . or pip install -e . for editable mode (if you want to make changes to mamba code).Then one should be able to load any iteration using the following snippet (say iteration 10,000):
from mamba_ssm.models.mixer_seq_simple import MambaLMHeadModel
model = MambaLMHeadModel.from_pretrained("Zyphra/Mamba-370M", iteration=10_000, device="cuda")
If iteration is not specified, then the model from the root of the repository is loaded, which is the final iteration (610,351).
The model was trained using "EleutherAI/gpt-neox-20b" tokenizer.
Here is a snippet for text generation:
import transformers, torch
tokenizer = transformers.AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")
inp_ids = torch.as_tensor([tokenizer.encode("Hello! How are you?")]).to("cuda")
out_ids = model.generate(inp_ids, max_length=100)
print(tokenizer.decode(out_ids[0]))