Downloads · 30 days
0
PAkshayV/LunarLander-v2
LunarLander-v2 is a reinforcement learning model from PAkshayV. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
This is a trained Proximal Policy Optimization (PPO) agent playing LunarLander-v2.
Downloads · 30 days
0
Access
Public
Updated Sep 7, 2026
Repo size
43.3 KB
Likes
0
Public
Click a slice to open those files.
.pt43.3 KB · 83%
From the Hugging Face model README
This is a trained Proximal Policy Optimization (PPO) agent playing LunarLander-v2.
This project was completed as part of the Hugging Face Deep Reinforcement Learning Course, Unit 8 - Part 1.
The PPO agent was implemented from scratch using PyTorch and Gymnasium, following the PPO implementation and concepts covered in the course.
The agent was trained on LunarLander-v2.
The goal is to learn a policy that controls a lunar lander and successfully lands it on the landing pad while maximizing the cumulative reward.
The environment provides an 8-dimensional observation describing:
The action space contains four discrete actions:
| Action | Description |
|---|---|
| 0 | Do nothing |
| 1 | Fire left orientation engine |
| 2 | Fire main engine |
| 3 | Fire right orientation engine |
The agent uses Proximal Policy Optimization (PPO).
PPO is an on-policy policy-gradient reinforcement learning algorithm that improves the policy while limiting excessively large policy updates.
The implementation includes:
The probability ratio between the current and old policies is:
r_t(theta) = pi_theta(a_t | s_t) / pi_theta_old(a_t | s_t)
The PPO clipped objective is:
L_CLIP = E[min(r_t A_t, clip(r_t, 1-epsilon, 1+epsilon) A_t)]
The clipping coefficient used is:
epsilon = 0.2
Therefore, the clipping range is:
[0.8, 1.2]
Clipping prevents the policy from making excessively large updates.
When the ratio is within the clipping range, the policy can be updated normally.
When the ratio moves outside the range in a direction that would make the policy update excessively large, the clipped objective limits the update.
The implementation uses Generalized Advantage Estimation (GAE).
The temporal-difference error is:
delta_t = r_t + gamma * V(s_t+1) - V(s_t)
The advantage is estimated recursively using:
A_t = delta_t + gamma * lambda * A_t+1
The implementation uses:
GAE provides a balance between bias and variance when estimating advantages.
The PPO agent uses an Actor-Critic architecture.
The Actor receives the environment observation and produces action logits.
Architecture:
Input -> Linear(64) -> Tanh -> Linear(64) -> Tanh -> Linear(4)
The output is used to create a categorical probability distribution over the four possible actions.
The Critic estimates the value of the current state.
Architecture:
Input -> Linear(64) -> Tanh -> Linear(64) -> Tanh -> Linear(1)
| Parameter | Value |
|---|---|
| Environment | LunarLander-v3 |
| Algorithm | PPO |
| Framework | PyTorch |
| Environment library | Gymnasium |
| Total timesteps | 100,000 |
| Learning rate | 0.00025 |
| Number of environments | 8 |
| Steps per rollout | 128 |
| Minibatches | 4 |
| Update epochs | 4 |
| Discount factor | 0.99 |
| GAE Lambda | 0.95 |
| PPO clip coefficient | 0.2 |
| Entropy coefficient | 0.01 |
| Value function coefficient | 0.5 |
| Maximum gradient norm | 0.5 |
| Advantage normalization | Enabled |
| GAE | Enabled |
| Learning-rate annealing | Enabled |
| Clipped value loss | Enabled |
| Random seed | 1 |
The trained agent was evaluated for 10 episodes.
The evaluation results are stored in evaluation.txt.
The file contains:
The evaluation result shown at the top of this model card should be replaced with the actual mean reward from evaluation.txt.
The trained model is stored as model.pt.
The model contains the PyTorch state dictionary of the trained Actor-Critic agent.
The same Agent architecture must be recreated before loading the weights.
Example:
import torch
agent.load_state_dict(
torch.load(
"model.pt",
map_location="cpu"
)
)
agent.eval()
The environment can be created using:
import gymnasium as gym
env = gym.make("LunarLander-v2")
observation, info = env.reset()
model.pt - trained PPO Actor-Critic modelhyperparameters.txt - PPO training hyperparametersevaluation.txt - evaluation resultsREADME.md - model cardThis project was completed as part of the:
Hugging Face Deep Reinforcement Learning Course
Unit 8 - Part 1: Proximal Policy Optimization (PPO)
The project demonstrates the implementation of PPO from scratch and its application to the LunarLander environment.
The main learning objectives include:
This model was developed as an educational project for the Hugging Face Deep Reinforcement Learning Course.
Performance may vary depending on: