Downloads · 30 days
42
0% of all-time downloads
VMware/open-llama-13b-open-instruct
open-llama-13b-open-instruct is a text generation model from VMware. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-sa-3.0.
Instruction-tuned version of the fully trained Open LLama 13B model. The model is open for <bCOMMERCIAL USE</b. <br
Downloads · 30 days
42
0% of all-time downloads
All-time downloads
9K
Public
Parameters
13B
52.1 GB on disk
Likes
18
Public
Click a slice to open those files.
.bin26 GB · 50%
How the weights are stored.
F1613B · 100%
From the Hugging Face model README
Instruction-tuned version of the fully trained Open LLama 13B model. The model is open for <b>COMMERCIAL USE</b>. <br>
<b> NOTE </b> : The model was trained using the Alpaca prompt template
<b> NOTE </b> : Fast tokenizer results in incorrect encoding, set the use_fast = False parameter, when instantiating the tokenizer
<b> NOTE </b> : The model might struggle with code as the tokenizer merges multiple spaces
import os
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = 'VMware/open-llama-13b-open-instruct'
tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=False)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, device_map='sequential')
prompt_template = "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Response:"
prompt = 'Explain in simple terms how the attention mechanism of a transformer model works'
inputt = prompt_template.format(instruction= prompt)
input_ids = tokenizer(inputt, return_tensors="pt").input_ids.to("cuda")
output1 = model.generate(input_ids, max_length=512)
input_length = input_ids.shape[1]
output1 = output1[:, input_length:]
output = tokenizer.decode(output1[0])
print(output)
The finetuning scripts will be available in our RAIL Github Repository
<B>TODO</B>