Downloads · 30 days
1.1K
9% of all-time downloads
duyntnet/stable-code-instruct-3b-imatrix-GGUF
stable-code-instruct-3b-imatrix-GGUF is a text generation model from duyntnet. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Quantizations of https://huggingface.co/stabilityai/stable-code-instruct-3b
Downloads · 30 days
1.1K
9% of all-time downloads
All-time downloads
12.5K
Public
Repo size
35.4 GB
Likes
0
Public
Click a slice to open those files.
.gguf35.4 GB · 100%
From the Hugging Face model README
Quantizations of https://huggingface.co/stabilityai/stable-code-instruct-3b
Here's how you can run the model use the model:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("stabilityai/stable-code-instruct-3b", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("stabilityai/stable-code-instruct-3b", torch_dtype=torch.bfloat16, trust_remote_code=True)
model.eval()
model = model.cuda()
messages = [
{
"role": "system",
"content": "You are a helpful and polite assistant",
},
{
"role": "user",
"content": "Write a simple website in HTML. When a user clicks the button, it shows a random joke from a list of 4 jokes."
},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer([prompt], return_tensors="pt").to(model.device)
tokens = model.generate(
**inputs,
max_new_tokens=1024,
temperature=0.5,
top_p=0.95,
top_k=100,
do_sample=True,
use_cache=True
)
output = tokenizer.batch_decode(tokens[:, inputs.input_ids.shape[-1]:], skip_special_tokens=False)[0]