Downloads · 30 days
14
21% of all-time downloads
codenopro/lab22-dpo-vn
lab22-dpo-vn is a text generation model from codenopro. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A LoRA adapter that DPO-aligns Qwen2.5-3B for Vietnamese instruction following. Pipeline: SFT-mini → DPO (the DPO LoRA continues training the SFT LoRA, so this single adapter is the final SFT+DPO model). Built for Day…
Downloads · 30 days
14
21% of all-time downloads
All-time downloads
68
Public
Repo size
131 MB
Likes
0
Public
Click a slice to open those files.
.safetensors120 MB · 91%
From the Hugging Face model README
A LoRA adapter that DPO-aligns Qwen2.5-3B for Vietnamese instruction following. Pipeline: SFT-mini → DPO (the DPO LoRA continues training the SFT LoRA, so this single adapter is the final SFT+DPO model). Built for Day 22 / Track 3 (DPO/ORPO Alignment) of the VinUni AICB program on a free Colab T4.
This adapter was produced by (1) a small supervised fine-tune of unsloth/Qwen2.5-3B-bnb-4bit on a 1,000-sample Vietnamese Alpaca slice, then (2) Direct Preference Optimization (TRL DPOTrainer, β=0.1) on 2,000 binarized UltraFeedback preference pairs. It is a teaching-scale run: DPO shifted the implicit-reward margin but produced only marginal, noisy behavior changes (see Evaluation). It is not a production or safety-aligned model.
vi)unsloth/Qwen2.5-3B-bnb-4bitVietnamese instruction following / chat using the Qwen2.5 ChatML template. Intended for research and education on the SFT→DPO alignment pipeline.
A starting point for further preference-tuning experiments (e.g. β-sweeps, ORPO/SimPO comparisons) or as a worked example of LoRA DPO on a 4-bit base.
Not for production, user-facing assistants, factual/medical/legal advice, or any safety-sensitive setting. The model is not safety-aligned (see Limitations) and degenerates on a small SFT base.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Do not deploy it where unsafe or low-quality completions could cause harm; treat outputs as illustrative, not reliable.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "Qwen/Qwen2.5-3B" # or load unsloth/Qwen2.5-3B-bnb-4bit in 4-bit
tok = AutoTokenizer.from_pretrained("codenopro/lab22-dpo-vn")
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.float16, device_map="auto")
model = PeftModel.from_pretrained(model, "codenopro/lab22-dpo-vn")
msgs = [{"role": "user", "content": "Giải thích ngắn gọn cách thuật toán quicksort hoạt động."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
5CD-AI/Vietnamese-alpaca-gpt4-gg-translated — 1,000-sample slice, 1 epoch.argilla/ultrafeedback-binarized-preferences-cleaned — 2,000 (prompt, chosen, rejected) pairs, 1 epoch.Two LoRA phases on the same adapter: SFT (NB1), then DPO (NB3) with the reference model auto-derived from the PEFT base (adapter disabled for the reference forward pass — no second copy of weights).
Examples formatted to Qwen2.5 ChatML (<|im_start|> / <|im_end|>) via tokenizer.apply_chat_template. Max sequence length 512, max prompt length 256.
r=16, lora_alpha=32, dropout 0, target q,k,v,o,gate,up,down_projbeta=0.1, loss_type=sigmoid, lr 5e-7, 250 steps (1 epoch), effective batch 8DPO took ≈ 40 min on a free T4 — slower than usual because T4 (compute 7.5) cannot run xformers' grouped-query-attention backward, so a PyTorch SDPA math-backend attention fallback was used. Final DPO loss 0.7719; end reward gap (chosen − rejected) ≈ 0.14 (last step) / ≈ 0.20 (last-5 mean). Adapter weights only (LoRA), not full model.
8 held-out Vietnamese prompts (4 helpfulness + 4 safety) — data/eval/side_by_side.jsonl in the lab repo.
Disaggregated by prompt category: helpfulness vs safety.
Manual win/loss/tie between SFT-only and SFT+DPO on full generations (no API judge).
| Metric | Result |
|---|---|
| SFT+DPO vs SFT-only | 1 win / 1 loss / 6 ties |
| Helpfulness | 0W / 1L / 3T |
| Safety | 1W / 0L / 3T |
| Final DPO loss | 0.7719 |
| End reward gap | ≈ 0.14 |
DPO's one clear win is a cleaner safety refusal (underage-alcohol prompt); its one loss is a repetition collapse on a quicksort explanation where SFT-only stayed coherent. Net effect is essentially a wash — consistent with the small base and short training.
The reward gap rose steadily while the qualitative win-rate stayed flat (1–1–6). This is the classic DPO caveat: optimizing the implicit-reward margin (chosen − rejected) does not guarantee better generations, and can even introduce new failure modes (more repetition) on a weak base.
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Qwen2.5-3B decoder-only causal LM with a LoRA adapter; objective is the DPO sigmoid loss over preference pairs (with an SFT cross-entropy warm-up phase).
Single NVIDIA Tesla T4 (16 GB), 4-bit quantized base + LoRA.
Unsloth 2026.4.8 · TRL · PEFT 0.19.1 · Transformers 5.5.0 · PyTorch 2.10.0+cu128 (CUDA 12.8) · bitsandbytes.
Day 22 · Track 3 · VinUni AICB — DPO/ORPO Alignment lab. Base model: Qwen2.5-3B (Apache-2.0). Trained with Unsloth + TRL.
BibTeX:
N/A (course lab; no associated paper)
APA:
codenopro. (2026). lab22-dpo-vn: SFT→DPO LoRA adapter for Qwen2.5-3B (Vietnamese). Hugging Face. https://huggingface.co/codenopro/lab22-dpo-vn
chosen − rejected implicit reward; the headline DPO diagnostic.See the lab repository for the full pipeline, reward-curve plot, side-by-side comparison, and reflection: https://github.com/anhkiet75/Day22-Track3-DPO-Alignment-Lab
codenopro
Via the Hugging Face repository: https://huggingface.co/codenopro/lab22-dpo-vn