Downloads · 30 days
24
4% of all-time downloads
monuminu/indo-instruct-llama2-13b
indo-instruct-llama2-13b is a text generation model from monuminu. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
indo-instruct-llama2-32kmodel card Model Details Developed by: monuminu Backbone Model: LLaMA-2 Language(s): English Library: HuggingFace Transformers License: Fine-tuned checkpoints is licensed under the Non-Commerci…
Downloads · 30 days
24
4% of all-time downloads
All-time downloads
558
Public
Parameters
13B
26 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors26 GB · 100%
How the weights are stored.
F1613B · 100%
From the Hugging Face model README
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
tokenizer = AutoTokenizer.from_pretrained("monuminu/indo-instruct-llama2-13b")
model = AutoModelForCausalLM.from_pretrained(
"monuminu/indo-instruct-llama2-32k",
device_map="auto",
torch_dtype=torch.float16,
load_in_8bit=True,
)
prompt = "### User:\nThomas is healthy, but he has to go to the hospital. What could be the reasons?\n\n### Assistant:\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
del inputs["token_type_ids"]
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
output = model.generate(**inputs, streamer=streamer, use_cache=True, max_new_tokens=float('inf'))
output_text = tokenizer.decode(output[0], skip_special_tokens=True)