Downloads · 30 days
12
33% of all-time downloads
ketencrypt10n/ppo-lunar-lander
ppo-lunar-lander is a reinforcement learning model from ketencrypt10n. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
A Proximal Policy Optimization (PPO) agent trained on the LunarLander-v3 environment from Gymnasium, implemented from scratch using PyTorch.
Downloads · 30 days
12
33% of all-time downloads
All-time downloads
36
Public
Repo size
907 KB
Likes
0
Public
Click a slice to open those files.
.pth556 KB · 61%
From the Hugging Face model README
A Proximal Policy Optimization (PPO) agent trained on the LunarLander-v3 environment from Gymnasium, implemented from scratch using PyTorch.
This model implements a PPO agent with separate actor and critic networks, using Generalized Advantage Estimation (GAE) for stable and efficient training. The agent learns to land a lunar module safely between two flags on the moon's surface.
Actor Network (Policy):
Critic Network (Value Function):
{
"algorithm": "PPO",
"environment": "LunarLander-v3",
"total_episodes": 1000,
"n_steps": 128,
"batch_size": 16,
"epochs": 4,
"actor_hidden_size": 256,
"critic_hidden_size": 256,
"epsilon": 0.2,
"gae_lambda": 0.95,
"gamma": 0.99,
"actor_learning_rate": 0.0003,
"critic_learning_rate": 0.001,
"gradient_clip": 0.5,
"optimizer": "Adam"
}

The visualization shows the training performance over the last 100 episodes, including raw and smoothed rewards and episode durations.
pip install gymnasium torch numpy matplotlib
import torch
import torch.nn as nn
import gymnasium as gym
from huggingface_hub import hf_hub_download
# Define the actor network architecture (must match training)
class ActorNetwork(nn.Module):
def __init__(self, input_dims=8, output_dims=4, hidden_layer1=256, hidden_layer2=256):
super().__init__()
self.actor = nn.Sequential(
nn.Linear(input_dims, hidden_layer1),
nn.ReLU(),
nn.Linear(hidden_layer1, hidden_layer2),
nn.ReLU(),
nn.Linear(hidden_layer2, output_dims)
)
def forward(self, state):
dist = torch.distributions.Categorical(logits=self.actor(state))
return dist
# Download actor model
actor_path = hf_hub_download(
repo_id="ketencrypt10n/ppo-lunar-lander",
filename="actor_network.pth"
)
# Load model
device = 'cuda' if torch.cuda.is_available() else 'cpu'
actor = ActorNetwork(input_dims=8, output_dims=4).to(device)
actor.actor.load_state_dict(torch.load(actor_path, map_location=device))
actor.eval()
# Test the agent
env = gym.make("LunarLander-v3", render_mode="human")
for episode in range(30):
done = False
total_reward = 0
state, info = env.reset()
while not done:
state_tensor = torch.tensor(state, dtype=torch.float32, device=device)
with torch.no_grad():
dist = actor(state_tensor)
action = dist.sample()
state, reward, terminated, truncated, info = env.step(action.item())
total_reward += reward
done = terminated or truncated
print(f"Episode: {episode+1}, Total Reward: {total_reward:.2f}")
env.close()
The repository includes statistics_ppo.png which shows the complete training visualization with:
You can also load the data files for custom analysis:
import numpy as np
from huggingface_hub import hf_hub_download
# Download training history (last 100 episodes)
reward_path = hf_hub_download(repo_id="ketencrypt10n/ppo-lunar-lander", filename="reward_history_last100.npy")
duration_path = hf_hub_download(repo_id="ketencrypt10n/ppo-lunar-lander", filename="episode_durations_last100.npy")
rewards = np.load(reward_path)
durations = np.load(duration_path)
print(f"Average reward (last 100 episodes): {np.mean(rewards):.2f}")
print(f"Max reward (last 100 episodes): {np.max(rewards):.2f}")
print(f"Average duration (last 100 episodes): {np.mean(durations):.2f} steps")
This implementation was built from scratch using:
The PPO algorithm optimizes the clipped surrogate objective:
L^CLIP(θ) = E[min(r_t(θ)Â_t, clip(r_t(θ), 1-ε, 1+ε)Â_t)]
where:
r_t(θ) is the probability ratio between new and old policiesÂ_t is the generalized advantage estimateε is the clipping parameter (0.2)The Lunar Lander environment challenges the agent to:
Reward Structure:
@misc{ppo_lunar_lander,
author = {ketencrypt10n},
title = {PPO Agent for Lunar Lander},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/ketencrypt10n/ppo-lunar-lander}}
}
This model is released for educational and research purposes.