Downloads · 30 days
18
13% of all-time downloads
Jo1uck/mamba-11b-back
mamba-11b-back is a text generation model from Jo1uck. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains the transfromers compatible mamba. The checkpoints are untouched, but the full config.json and tokenizer are pushed to this repo.
Downloads · 30 days
18
13% of all-time downloads
All-time downloads
138
Public
Parameters
10.8B
43.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors43.1 GB · 100%
From the Hugging Face model README
This repository contains the transfromers compatible mamba. The checkpoints are untouched, but the full config.json and tokenizer are pushed to this repo.
You need to install transformers from main until transformers=4.39.0 is released.
pip install git+https://github.com/huggingface/transformers@main
We also recommend you to install both causal_conv_1d and mamba-ssm using:
pip install causal-conv1d>=1.2.0
pip install mamba-ssm
If any of these two is not installed, the "eager" implementation will be used. Otherwise the more optimised cuda kernels will be used.
You can use the classic generate API:
>>> from transformers import MambaConfig, MambaForCausalLM, AutoTokenizer
>>> import torch
>>> tokenizer = AutoTokenizer.from_pretrained("Jo1uck/mamba-11b-back")
>>> model = MambaForCausalLM.from_pretrained("Jo1uck/mamba-11b-back")
>>> input_ids = tokenizer("Hey how are you doing?", return_tensors="pt")["input_ids"]
>>> out = model.generate(input_ids, max_new_tokens=10)
>>> print(tokenizer.batch_decode(out))
["Hey how are you doing?\n\nI'm doing great.\n\nI"]
In order to finetune using the peft library, we recommend keeping the model in float32!
from datasets import load_dataset
from trl import SFTTrainer
from peft import LoraConfig
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments
tokenizer = AutoTokenizer.from_pretrained("Jo1uck/mamba-11b-back")
model = AutoModelForCausalLM.from_pretrained("Jo1uck/mamba-11b-back")
dataset = load_dataset("Abirate/english_quotes", split="train")
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=4,
logging_dir='./logs',
logging_steps=10,
learning_rate=2e-3
)
lora_config = LoraConfig(
r=8,
target_modules=["x_proj", "embeddings", "in_proj", "out_proj"],
task_type="CAUSAL_LM",
bias="none"
)
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
args=training_args,
peft_config=lora_config,
train_dataset=dataset,
dataset_text_field="quote",
)
trainer.train()