Downloads · 30 days
15
22% of all-time downloads
arif-butt/tinyllama-unsloth-gguf
tinyllama-unsloth-gguf is a machine learning model from arif-butt. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This is a GGUF quantized version of TinyLlama (1.1B parameters) fine-tuned using Unsloth (2x faster training, 50% less memory) with LoRA adapters. The model has been optimized for efficient CPU inference with minimal…
Downloads · 30 days
15
22% of all-time downloads
All-time downloads
67
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf1.2 GB · 100%
From the Hugging Face model README
This is a GGUF quantized version of TinyLlama (1.1B parameters) fine-tuned using Unsloth (2x faster training, 50% less memory) with LoRA adapters. The model has been optimized for efficient CPU inference with minimal memory footprint.
# Install llama-cpp-python
# pip install llama-cpp-python
from llama_cpp import Llama
# Load the model
model_path = "arif-butt/tinyllama-unsloth-gguf" # or local path to .gguf file
llm = Llama(
model_path=model_path,
n_ctx=2048, # Context length
n_threads=4, # Number of CPU threads
n_gpu_layers=0, # Set >0 for GPU offloading
verbose=False,
)
# Simple prompt
prompt = "Q: Name all the courses Arif butt teach?\nA:"
# Generate response
output = llm(
prompt,
max_tokens=100, # Maximum tokens to generate
temperature=0.2, # Lower = more deterministic
top_p=0.95, # Nucleus sampling
repeat_penalty=1.1, # Penalize repetition
stop=["Q:", "\nQ:"], # Stop sequences
)
print(output["choices"][0]["text"])
Option 2: Using Transformers with llama-cpp
from transformers import AutoTokenizer
from llama_cpp import Llama
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
# Load GGUF model
model_path = "arif-butt/tinyllama-unsloth-gguf"
llm = Llama(model_path=model_path, n_ctx=2048)
def generate_response(prompt, max_tokens=100, temperature=0.7):
"""Generate response using the GGUF model"""
output = llm(
prompt,
max_tokens=max_tokens,
temperature=temperature,
stop=["Q:", "\nQ:", "User:", "\nUser:"],
)
return output["choices"][0]["text"]
# Test
prompt = "Q: What is machine learning?\nA:"
response = generate_response(prompt)
print(f"Response: {response}")
Option 3: Using llama.cpp CLI
# Download the model
wget https://huggingface.co/arif-butt/tinyllama-unsloth-gguf/resolve/main/tinyllama-unsloth-q4_k_m.gguf
# Run inference
./main -m tinyllama-unsloth-q4_k_m.gguf \
-p "Q: Name all the courses Arif butt teach?\nA:" \
-n 100 \
-t 4 \
--temp 0.2 \
--top_p 0.95
LORA_R = 16 # Rank of LoRA matrices
LORA_ALPHA = 32 # Scaling factor (alpha/r = 2.0)
LORA_DROPOUT = 0.05 # Dropout for regularization
TARGET_MODULES = [ # Layers where LoRA is applied
"q_proj", # Query projection
"k_proj", # Key projection
"v_proj", # Value projection
"o_proj", # Output projection
"gate_proj", # Gate projection (MLP)
"up_proj", # Up projection (MLP)
"down_proj" # Down projection (MLP)
]