Downloads · 30 days
11
9% of all-time downloads
PaddySeahorse/tinyllama-4bit
tinyllama-4bit is a machine learning model from PaddySeahorse. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repo contains a 4-bit quantized version of the TinyLlama-1.1B base model, quantized in the summer of 2025. It is designed for ultra-lightweight inference, edge computing, and hardware benchmarking.
Downloads · 30 days
11
9% of all-time downloads
All-time downloads
125
Public
Parameters
1.1B
808 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors807 MB · 100%
How the weights are stored.
U8969M · 87%
From the Hugging Face model README
This repo contains a 4-bit quantized version of the TinyLlama-1.1B base model, quantized in the summer of 2025. It is designed for ultra-lightweight inference, edge computing, and hardware benchmarking.
bitsandbytes compatibletransformers, accelerate) without high VRAM/RAM overhead.You can easily load and run this model using Hugging Face transformers and bitsandbytes:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "PaddySeahorse/tinyllama-4bit"
# 4-bit configuration
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_quant_type="nf4"
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto"
)
# Inference
prompt = "The meaning of life is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Created in Summer 2025 as an exploration in LLM quantization.