Downloads · 30 days
0
cy0307/lm-dpo-alignment
lm-dpo-alignment is a text generation model from cy0307. Use it when you need the model to write or continue text. The card lists the license as mit.
Align an LLM to preferred answers with Direct Preference Optimization.
Downloads · 30 days
0
Access
Public
Updated Jun 28, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md4.4 KB · 72%
From the Hugging Face model README
Align an LLM to preferred answers with Direct Preference Optimization.
Status — documented recipe (placeholder). A production-grade pipeline from Ropedia Academy for an advanced, GPU-heavy task. Everything below — base model, objective, dataset, config, the exact evaluation — is specified; the weights / metrics / figures land here automatically when you run the notebook on a GPU (one click below). Try the trained models live in the Ropedia demos Space.
| Base model | An SFT'd LLM (e.g. Qwen2.5-0.5B-Instruct, 4-bit) |
| Task | preference alignment |
| Training objective | Direct Preference Optimization on chosen/rejected pairs (no reward model). |
| Track | LM · Language & multimodal |
| Built on | huggingface/trl |
| Notebook | |
| Compute / storage / time | GPU required — see the Compute · storage · time table in the notebook |
GPU-scale — the notebook ships a demo profile (free Colab T4) and a full profile, with an exact Compute · storage · time table. Hyperparameters (optimizer, steps, batch, LoRA rank, …) are in the training cell.
⏳ Pending — run the notebook on a GPU to fill this in. This lab reports preference accuracy (eval_rewards/accuracies) on a held-out split (see its Evaluate cell).
No weights are published yet. After a GPU run, load the checkpoint/adapter the notebook saves (it also has a ready inference cell). Base model: An SFT'd LLM (e.g. Qwen2.5-0.5B-Instruct, 4-bit).
HfApi().upload_folder(...)) — the checkpoint + metrics.json + figures replace this placeholder.metrics.json · [ ] add figures · [ ] swap in the real results cardNot yet trained — no numbers to report. The pipeline is GPU-heavy (see the compute table); on free Colab use the demo-scale settings. This is an educational, reproducible recipe, not a tuned production release.
Code: MIT (this repository). The base model (huggingface/trl) and dataset are each under their own licenses — check the upstream source before redistribution.
@misc{ropedia_academy,
title = {Ropedia Academy: an interactive course on embodied & spatial AI},
author = {Ropedia Academy},
year = {2026},
howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}
Method / original work: Rafailov et al., DPO, NeurIPS 2023.
Documented placeholder in the Ropedia Academy collection — train it on a GPU to publish the real model. Contributions welcome on GitHub.