Downloads · 30 days
0
DavidH2802/PPO-from-scratch
PPO-from-scratch is a reinforcement learning model from DavidH2802. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
A Proximal Policy Optimization (PPO) policy trained from scratch in PyTorch on the Isaac-Reach-Franka-v0 task using NVIDIA Isaac Lab with 4096 GPU-parallel environments.
Downloads · 30 days
0
Access
Public
Updated Apr 24, 2026
Repo size
2.6 MB
Likes
0
Public
Click a slice to open those files.
.gif2 MB · 76%
From the Hugging Face model README
A Proximal Policy Optimization (PPO) policy trained from scratch in PyTorch on the Isaac-Reach-Franka-v0 task using NVIDIA Isaac Lab with 4096 GPU-parallel environments.
GitHub Repository: DavidH2802/PPO-from-scratch
<p align="center"> <img src="franka_reach.gif" alt="Franka Reach Policy" width="480"/> </p>The model is a diagonal Gaussian policy (Actor) that controls a 7-DOF Franka Emika robot arm to reach a randomly spawned target position in 3D space. The policy outputs continuous joint-level actions.
| Parameter | Value |
|---|---|
| Task | Isaac-Reach-Franka-v0 |
| Parallel Envs | 4096 |
| Learning Rate | 3e-4 |
| Discount (γ) | 0.99 |
| GAE (λ) | 0.95 |
| Clip (ε) | 0.2 |
| Epochs per Update | 4 |
| Minibatch Size | 2048 |
| Horizon | 32 |
| Total Iterations | 500 |
| Total Env Steps | 65.5M |
| Training Time | ~48 minutes |
The agent starts with negative reward (arm far from target) and converges to positive reward (~0.03-0.05) as it learns to reach the target.
The checkpoint includes running mean and variance statistics for observation normalization. These must be restored at inference time — without them, the policy receives unnormalized inputs and will not perform correctly.
from huggingface_hub import hf_hub_download
checkpoint_path = hf_hub_download(
repo_id="DavidH2802/PPO-from-scratch",
filename="final_policy.pt",
)
Clone the full project for the model and environment code:
git clone https://github.com/DavidH2802/PPO-from-scratch.git
cd PPO-from-scratch
See the GitHub repository for complete setup instructions including Isaac Lab installation and the eval.py script for video recording.
The final_policy.pt file contains:
| Key | Description |
|---|---|
actor | Actor network state dict |
critic | Critic network state dict |
obs_rms_mean | Running mean for observation normalization |
obs_rms_var | Running variance for observation normalization |
@misc{habinski2026ppo,
author = {David Habinski},
title = {PPO from Scratch in PyTorch with Isaac Lab},
year = {2026},
publisher = {GitHub},
url = {https://github.com/DavidH2802/PPO-from-scratch}
}
MIT