Downloads · 30 days
49
0% of all-time downloads
RedHatAI/llama-2-7b-chat-marlin
llama-2-7b-chat-marlin is a text generation model from RedHatAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
Example of converting a GPTQ model to Marlin format for fast batched decoding with Marlin Kernels
Downloads · 30 days
49
0% of all-time downloads
All-time downloads
25.1K
Public
Parameters
1.1B
3.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.9 GB · 100%
How the weights are stored.
I32810M · 72%
From the Hugging Face model README
Example of converting a GPTQ model to Marlin format for fast batched decoding with Marlin Kernels
pip install torch
git clone https://github.com/IST-DASLab/marlin.git
cd marlin
pip install -e .
Convert the model from GPTQ to Marlin format. Note that this requires:
sym=truegroup_size=128desc_activations=falsepip install -U transformers accelerate auto-gptq optimum
Convert with the convert.py script in this repo:
python3 convert.py --model-id "TheBloke/Llama-2-7B-Chat-GPTQ" --save-path "./marlin-model" --do-generation
Load with the load.load_model utility from this repo and run inference as usual.
from load import load_model
from transformers import AutoTokenizer
# Load model from disk.
model_path = "./marlin-model"
model = load_model(model_path).to("cuda")
tokenizer = AutoTokenizer.from_pretrained(model_path)
# Generate text.
inputs = tokenizer("My favorite song is", return_tensors="pt")
inputs = {k: v.to("cuda") for k, v in inputs.items()}
outputs = model.generate(**inputs, max_new_tokens=50, do_sample=False)
print(tokenizer.batch_decode(outputs)[0])