Downloads · 30 days
0
mariam8/tinyllama-xsum-lora
tinyllama-xsum-lora is a machine learning model from mariam8. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
--- license: apache-2.0 basemodel: TinyLlama/TinyLlama-1.1B-Chat-v1.0 libraryname: peft tags: - lora - peft - summarization - text-generation - tinyllama datasets: - EdinburghNLP/xsum language: - en pipelinetag: summa…
Downloads · 30 days
0
Access
Public
Updated Aug 19, 2026
Repo size
9 MB
Likes
0
Public
Click a slice to open those files.
.safetensors9 MB · 71%
From the Hugging Face model README
license: apache-2.0 base_model: TinyLlama/TinyLlama-1.1B-Chat-v1.0 library_name: peft tags:
This is a LoRA (PEFT) adapter fine-tuned on top of TinyLlama/TinyLlama-1.1B-Chat-v1.0 for single-sentence news summarization, trained on a subset of the XSum dataset.
This is an adapter only — it must be loaded on top of the base model (see Usage below).
Evaluated on a held-out sample of 30 examples from the XSum validation split, comparing the base model (zero-shot) against the LoRA fine-tuned model.
| Metric | Before (base, zero-shot) | After (LoRA fine-tuned) | Improvement |
|---|---|---|---|
| ROUGE-1 | 0.138 | 0.238 | +72% |
| ROUGE-2 | — | 0.096 | — |
| ROUGE-L | — | 0.184 | — |
| ROUGE-Lsum | — | 0.191 | — |
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
base_model_name = "TinyLlama/TinyLlama-1.1B-Chat-v1.0"
adapter_name = "your-username/tinyllama-xsum-lora" # replace with your repo id
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
)
tokenizer = AutoTokenizer.from_pretrained(base_model_name)
base_model = AutoModelForCausalLM.from_pretrained(
base_model_name,
quantization_config=bnb_config,
torch_dtype=torch.float16,
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, adapter_name)
prompt = (
"<|system|>\nYou are a helpful assistant that summarizes news articles "
"in one short sentence.</s>\n"
"<|user|>\nSummarize the following article:\n{your_article_here}</s>\n"
"<|assistant|>\n"
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=60, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Educational / portfolio project demonstrating parameter-efficient fine-tuning (LoRA/PEFT) of a small LLM for a summarization task under limited compute (free-tier Colab GPU).