Downloads · 30 days
7
8% of all-time downloads
arunma/monty
monty is a text generation model from arunma. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A small LoRA adapter that gives Qwen2.5-0.5B-Instruct an opinionated, slightly-grumpy "Monty" voice. Trained as the Phase A milestone of learn-you-an-sft — a from-scratch tour of supervised fine-tuning.
Downloads · 30 days
7
8% of all-time downloads
All-time downloads
90
Public
Repo size
2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors29.5 MB · 72%
From the Hugging Face model README
A small LoRA adapter that gives Qwen2.5-0.5B-Instruct an opinionated, slightly-grumpy "Monty" voice. Trained as the Phase A milestone of learn-you-an-sft — a from-scratch tour of supervised fine-tuning.
The base model weights are not distributed here. This repo ships only the LoRA adapter (adapter_model.safetensors, ~tens of MB) plus tokenizer config. You merge it onto the base at load time.
Qwen/Qwen2.5-0.5B-Instructruns/sft_v1_trl/train.pyEducational. Demonstrates how a few thousand persona-shaped Q&A pairs can shift a small instruction-tuned model's voice and disposition without touching factual knowledge.
Treat outputs as drafts, not facts. If you fork this for your own persona, plan on a separate evaluation pass.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "Qwen/Qwen2.5-0.5B-Instruct"
adapter_id = "arunma/monty"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
base = AutoModelForCausalLM.from_pretrained(base_id, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, adapter_id)
model.eval()
messages = [
{"role": "system", "content": "You are Monty."},
{"role": "user", "content": "Should I learn Rust?"},
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
out = model.generate(inputs, max_new_tokens=200, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
lid.176, English only) → MinHash near-dedup → 95/5 train/val split.data/ pipelineSFTTrainer with assistant_only_loss=True (loss masked to assistant tokens only via the Qwen chat template).q_proj, k_proj, v_proj, o_proj).| Setting | Value |
|---|---|
LoRA rank r | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Epochs | 3 |
| Per-device batch size | 8 |
| Max sequence length | 1024 |
| Learning rate | 2e-4 |
| LR scheduler | cosine |
| Warmup ratio | 0.03 |
| Weight decay | 0.0 |
| Optimizer | AdamW (default) |
| Precision | bfloat16 |
| Gradient checkpointing | enabled |
Currently uses only training-loop signals (train loss, eval loss on the 273-pair val split, mean token accuracy). A judge-based persona-fidelity eval (Lesson 8 of the parent project) is planned but not yet attached to this checkpoint.