Downloads · 30 days
100
8% of all-time downloads
RedHatAI/Meta-Llama-3-70B-Instruct-quantized.w8a8
Meta-Llama-3-70B-Instruct-quantized.w8a8 is a text generation model from RedHatAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.
- Model Architecture: Meta-Llama-3 - Input: Text - Output: Text - Model Optimizations: - Activation quantization: INT8 - Weight quantization: INT8 - Intended Use Cases: Intended for commercial and research use in Engl…
Downloads · 30 days
100
8% of all-time downloads
All-time downloads
1.3K
Public
Parameters
70.6B
72.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors72.7 GB · 100%
How the weights are stored.
I868.5B · 97%
From the Hugging Face model README
Quantized version of Meta-Llama-3-70B-Instruct. It achieves an average score of 79.18 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 79.18.
This model was obtained by quantizing the weights of Meta-Llama-3-70B-Instruct to INT8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x). Weight quantization also reduces disk size requirements by approximately 50%.
Only weights and activations of the linear operators within transformers blocks are quantized. Weights are quantized with a symmetric static per-channel scheme, where a fixed linear scaling factor is applied between INT8 and floating point representations for each output channel dimension. Activations are quantized with a symmetric dynamic per-token scheme, computing a linear scaling factor at runtime for each token between INT8 and floating point representations. The GPTQ algorithm is applied for quantization, as implemented in the llm-compressor library. GPTQ used a 10% damping factor and 256 sequences taken from Neural Magic's LLM compression calibration dataset.
This model can be deployed efficiently using the vLLM backend, as shown in the example below (using 2 GPUs).
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
model_id = "neuralmagic/Meta-Llama-3-70B-Instruct-quantized.w8a8"
number_gpus = 2
sampling_params = SamplingParams(temperature=0.6, top_p=0.9, max_tokens=256)
tokenizer = AutoTokenizer.from_pretrained(model_id)
messages = [
{"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
{"role": "user", "content": "Who are you?"},
]
prompts = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
llm = LLM(model=model_id, tensor_parallel_size=number_gpus)
outputs = llm.generate(prompts, sampling_params)
generated_text = outputs[0].outputs[0].text
print(generated_text)
vLLM aslo supports OpenAI-compatible serving. See the documentation for more details.
The following example contemplates how the model can be deployed in Transformers using the generate() function.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "neuralmagic/Meta-Llama-3-70B-Instruct-quantized.w8a8"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [
{"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
{"role": "user", "content": "Who are you?"},
]
input_ids = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
terminators = [
tokenizer.eos_token_id,
tokenizer.convert_tokens_to_ids("<|eot_id|>")
]
outputs = model.generate(
input_ids,
max_new_tokens=256,
eos_token_id=terminators,
do_sample=True,
temperature=0.6,
top_p=0.9,
)
response = outputs[0][input_ids.shape[-1]:]
print(tokenizer.decode(response, skip_special_tokens=True))
This model was created by using the llm-compressor library as presented in the code snipet below.
from transformers import AutoTokenizer
from datasets import load_dataset
from llmcompressor.transformers import SparseAutoModelForCausalLM, oneshot
from llmcompressor.modifiers.quantization import GPTQModifier
model_id = "meta-llama/Meta-Llama-3-70B-Instruct"
num_samples = 256
max_seq_len = 8192
tokenizer = AutoTokenizer.from_pretrained(model_id)
def preprocess_fn(example):
return {"text": tokenizer.apply_chat_template(example["messages"], add_generation_prompt=False, tokenize=False)}
ds = load_dataset("neuralmagic/LLM_compression_calibration", split="train")
ds = ds.shuffle().select(range(num_samples))
ds = ds.map(preprocess_fn)
recipe = GPTQModifier(
targets="Linear",
scheme="W8A8",
ignore=["lm_head"],
dampening_frac=0.1,
)
model = SparseAutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
trust_remote_code=True,
)
oneshot(
model=model,
dataset=ds,
recipe=recipe,
max_seq_length=max_seq_len,
num_calibration_samples=num_samples,
)
model.save_pretrained("Meta-Llama-3-70B-Instruct-quantized.w8a8")
The model was evaluated on the OpenLLM leaderboard tasks (version 1) with the lm-evaluation-harness (commit 383bbd54bc621086e05aa1b030d8d4d5635b25e6) and the vLLM engine, using the following command (using 2 GPUs):
lm_eval \
--model vllm \
--model_args pretrained="neuralmagic/Meta-Llama-3-70B-Instruct-quantized.w8a8",tensor_parallel_size=2,dtype=auto,gpu_memory_utilization=0.4,add_bos_token=True,max_model_len=4096 \
--tasks openllm \
--batch_size auto