Downloads · 30 days
51
27% of all-time downloads
VHRamirez/victor-ramirez-7b-lora
victor-ramirez-7b-lora is a machine learning model from VHRamirez. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
A ~10 MB LoRA adapter that teaches a Qwen2.5-7B base model to answer professional / technical questions in Victor Ramirez's register. It learns phrasing, tone and structure only -- a companion RAG system supplies ever…
Downloads · 30 days
51
27% of all-time downloads
All-time downloads
187
Public
Repo size
324 MB
Likes
1
Public
Click a slice to open those files.
.json11.4 MB · 53%
From the Hugging Face model README
A ~10 MB LoRA adapter that teaches a Qwen2.5-7B base model to answer professional / technical questions in Victor Ramirez's register. It learns phrasing, tone and structure only -- a companion RAG system supplies every fact at inference. The goal: replace a 70B model in the AI-Vic chatbot with a 7B base + this adapter at a fraction of the cost, without losing answer quality.
High-scoring (instruction, context, output) triples from the AI-Vic evaluation loop,
scored by an LLM-as-judge (@cf/meta/llama-3.1-8b-instruct-fast) and kept only when:
Exported by .github/workflows/export-training-data.yml.
| Examples | 66 (52 train / 14 eval, seeded 80/20 split) |
| Format | instruction (user query) + context (top-k RAG chunks) -> output (target reply) |
| Source | https://gist.githubusercontent.com/vhr1975/7684de95fea6870bb3e9360c3703205b/raw/training-data.jsonl |
QLoRA: the base model is loaded in 4-bit (NF4) and frozen; only the adapter weights
train. Loss is computed on the reply tokens only (prompt tokens masked to -100), so
the model learns to answer, not to echo the prompt.
| Parameter | Value | What it is | Why this value |
|---|---|---|---|
| LoRA rank (r) | 8 | Size of the low-rank update added to each target weight | Small dataset (66 rows) -- enough capacity for style, not so much it memorises noise |
| LoRA alpha | 16 | Scales how strongly the adapter is applied (effective LR is proportional to alpha/r) | Kept at 2x r, the common ratio |
| Target modules | q_proj, v_proj | Which weight matrices get an adapter | Query + value projections carry most of the "voice"; cheaper than adapting every layer |
| Dropout | 0.05 | Fraction of adapter activations dropped each step | Light regularisation against overfitting a tiny dataset |
| Max sequence length | 1024 | Longest (prompt + reply) kept; context is left-truncated to fit | Covers the RAG context blocks; longer = more VRAM and slower steps |
| Per-device batch | 1 | Examples per forward pass on the GPU | What fits a free-tier T4 in 4-bit |
| Gradient accumulation | 4 | Forward passes before a weight update | Effective batch = 4; smooths gradients without more VRAM |
| Learning rate | 0.0002 | Step size for the adapter weights | Standard for LoRA -- higher than full fine-tuning, since the frozen base has nothing to forget |
| LR schedule | SchedulerType.COSINE | How the LR changes over training | Warm up, then decay toward 0 for a stable finish |
| Warmup steps | 3 | Steps to ramp the LR from 0 to full | Lets the optimizer settle before large updates |
| Epochs | 5 | Passes over the training set | About 65 optimizer steps total (13/epoch) -- enough to pick up register, few enough to not overfit |
| Weight decay | 0.01 | L2 penalty on weights | Light overfitting guard |
| Precision / optimizer | fp16 + paged_adamw_8bit | Mixed-precision training, 8-bit optimizer states | Fits the T4; paging avoids OOM spikes |
The best checkpoint (lowest eval loss, not the final epoch) is reloaded before saving. Best eval loss this run: 1.046. Trained on a Colab free-tier Tesla T4 in roughly 10-15 minutes.
The adapter is not scored in isolation -- it is evaluated inside AI-Vic against the current production path (retrieval + a larger model):
Step 8 of the training notebook runs a base-vs-tuned judge comparison on the held-out
rows and prints a pass/fail (tuned must match or beat base on both axes). Production A/B
numbers live in ai-vic-chatbot/docs/evaluation.md; treat any figure not from a real
run as pending.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
BASE = "Qwen/Qwen2.5-7B-Instruct"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16)
base = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(base, "VHRamirez/victor-ramirez-7b-lora")
tok = AutoTokenizer.from_pretrained(BASE)
# The adapter expects the ChatML prompt format it was trained on:
prompt = tok.apply_chat_template(
[{"role": "user", "content": "<question>\n\nContext: <retrieved chunks>"}],
add_generation_prompt=True, tokenize=False,
)
ids = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
print(tok.decode(
model.generate(**ids, max_new_tokens=256, do_sample=True, temperature=0.7)[0][ids["input_ids"].shape[1]:],
skip_special_tokens=True,
))
For a single standalone model, model.merge_and_unload() then save_pretrained / push.
For a Hugging Face Inference Endpoint, point it at the base model and attach this
adapter. Cloudflare Workers AI does not load HF adapters automatically -- it needs
its own LoRA upload against a supported base model (see the Workers AI LoRA docs).
Qwen/Qwen2.5-7B-Instruct to use it.@misc{ramirez2026victor7blora,
title = {Victor Ramirez 7B LoRA Adapter},
author = {Ramirez, Victor},
year = {2026},
howpublished = {\url{https://huggingface.co/VHRamirez/victor-ramirez-7b-lora}}
}
Training notebook: docs/phase-4b-colab-template.ipynb · Repo: ramirez-ai-labs/ai-vic-chatbot