Downloads · 30 days
68
4% of all-time downloads
VishaalY/CodeLlama-70b-instruct-neuron
CodeLlama-70b-instruct-neuron is a text generation model from VishaalY. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
This repo shows how you can utilize AWS-designed silicon to run inference on Codellama-70B-Instruct-hf! I ran this model on HumanEval locally and was getting 22.58237868454958 tokens per second running on an inf2.48xl…
Downloads · 30 days
68
4% of all-time downloads
All-time downloads
1.6K
Public
Repo size
276 GB
Likes
1
Public
Click a slice to open those files.
.weight276 GB · 100%
From the Hugging Face model README
This repo shows how you can utilize AWS-designed silicon to run inference on Codellama-70B-Instruct-hf! I ran this model on HumanEval locally and was getting 22.58237868454958 tokens per second running on an inf2.48xlarge.
The example below shows a single sample.
def string_to_md5(text):
"""
Given a string 'text', return its md5 hash equivalent string.
If 'text' is an empty string, return None.
>>> string_to_md5('Hello world') == '3e25960a79dbc69b674cd4ec67a72c62'
"""
from hashlib import md5
if not isinstance(text, str) or text == '':
return None
return ''.join([i for i in md5(bytes(text.encode('ascii'))).hexdigest()])
if __name__ == '__main__':
import doctest
doctest.testmod()
Launch an inf2.48xlarge instance using Amazon EC2. Use the HuggingFace Neuron DLAMI.
Use the commands below to install the following packages or create a bash script. You can run the following commands in your terminal.
sudo apt-get update -y \
&& sudo apt-get install -y --no-install-recommends \
aws-neuronx-dkms=2.15.9.0 \
aws-neuronx-collectives=2.19.7.0-530fb3064 \
aws-neuronx-runtime-lib=2.19.5.0-97e2d271b \
aws-neuronx-tools=2.16.1.0
pip3 install --upgrade \
neuronx-cc==2.12.54.0 \
torch-neuronx==1.13.1.1.13.0 \
transformers-neuronx==0.9.474 \
--extra-index-url=https://pip.repos.neuron.amazonaws.com
git lfs clone https://huggingface.co/VishaalY/CodeLlama-70b-instruct-neuron
import torch
from transformers_neuronx.module import save_pretrained_split
from transformers import LlamaForCausalLM
from transformers_neuronx.config import NeuronConfig
from transformers_neuronx import constants
from sentencepiece import SentencePieceProcessor
import time
from transformers import AutoTokenizer
from transformers_neuronx.llama.model import LlamaForSampling
import os
print("construct a tokenizer and encode prompt text")
tokenizer = AutoTokenizer.from_pretrained('codellama/CodeLlama-70b-hf')
# ----------------------------------------------------------------------------------------
print("Load from Neuron Artifacts")
neuron_model = LlamaForSampling.from_pretrained('./CodeLlama-70b-Instruct-hf/', batch_size=1, tp_degree=24, amp='f16')
neuron_model.load('./CodeLlama-70b-Instruct-hf/') # Load the compiled Neuron artifacts
neuron_model.to_neuron() # will skip compile
# ------------------------------------------------------s---------------------------------------------------------
while(True):
prompt = input("User: ")
input_ids = tokenizer.encode(prompt, return_tensors="pt")
with torch.inference_mode():
start = time.time()
generated_sequences = neuron_model.sample(input_ids, sequence_length=2048, temperature=0.1)
elapsed = time.time() - start
generated_sequences = [tokenizer.decode(seq) for seq in generated_sequences]
print(f'generated sequences {generated_sequences} in {elapsed} seconds')
print(generated_sequences[0])
if (input("Continue?") == "N"):
break
to deploy onto SageMaker follow these instructions and change the model identifiers to this repo.
input_shapes
{
"batch_size": 1,
"sequence_length": 2048,
}
compiler_args
{
"auto_cast_type": "bf16",
"num_cores": 24,
}