Downloads · 30 days
0
cy0307/lm-rlhf-ppo
lm-rlhf-ppo is a text generation model from cy0307. Use it when you need the model to write or continue text. The card lists the license as mit.
Optimize an LLM with PPO against a reward model — the classic RLHF recipe.
Downloads · 30 days
0
Access
Public
Updated Jun 28, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md4.4 KB · 73%
From the Hugging Face model README
Optimize an LLM with PPO against a reward model — the classic RLHF recipe.
Status — documented recipe (placeholder). A production-grade pipeline from Ropedia Academy for an advanced, GPU-heavy task. Everything below — base model, objective, dataset, config, the exact evaluation — is specified; the weights / metrics / figures land here automatically when you run the notebook on a GPU (one click below). Try the trained models live in the Ropedia demos Space.
| Base model | An SFT LLM + a reward model (demo: GPT-2 sentiment) |
| Task | RL fine-tuning against a reward model |
| Training objective | PPO — maximize reward-model score with a KL penalty to the reference. |
| Track | LM · Language & multimodal |
| Built on | huggingface/trl (PPOTrainer) |
| Notebook | |
| Compute / storage / time | GPU required — see the Compute · storage · time table in the notebook |
GPU-scale — the notebook ships a demo profile (free Colab T4) and a full profile, with an exact Compute · storage · time table. Hyperparameters (optimizer, steps, batch, LoRA rank, …) are in the training cell.
⏳ Pending — run the notebook on a GPU to fill this in. This lab reports mean reward (+ KL-to-reference) on a held-out split (see its Evaluate cell).
No weights are published yet. After a GPU run, load the checkpoint/adapter the notebook saves (it also has a ready inference cell). Base model: An SFT LLM + a reward model (demo: GPT-2 sentiment).
HfApi().upload_folder(...)) — the checkpoint + metrics.json + figures replace this placeholder.metrics.json · [ ] add figures · [ ] swap in the real results cardNot yet trained — no numbers to report. The pipeline is GPU-heavy (see the compute table); on free Colab use the demo-scale settings. This is an educational, reproducible recipe, not a tuned production release.
Code: MIT (this repository). The base model (huggingface/trl (PPOTrainer)) and dataset are each under their own licenses — check the upstream source before redistribution.
@misc{ropedia_academy,
title = {Ropedia Academy: an interactive course on embodied & spatial AI},
author = {Ropedia Academy},
year = {2026},
howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}
Method / original work: Ouyang et al., InstructGPT, 2022; Schulman et al., PPO, 2017.
Documented placeholder in the Ropedia Academy collection — train it on a GPU to publish the real model. Contributions welcome on GitHub.