Downloads · 30 days
0
viswa752/LunarLander-v2
LunarLander-v2 is a reinforcement learning model from viswa752. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
This repository contains my Proximal Policy Optimization (PPO) agent trained using PyTorch and Gymnasium on the LunarLander-v2 environment.
Downloads · 30 days
0
Access
Public
Updated Sep 14, 2026
Repo size
43.3 KB
Likes
0
Public
Click a slice to open those files.
.pt43.3 KB · 91%
From the Hugging Face model README
This repository contains my Proximal Policy Optimization (PPO) agent trained using PyTorch and Gymnasium on the LunarLander-v2 environment.
The agent uses Proximal Policy Optimization (PPO).
The implementation includes:
| Parameter | Value |
|---|---|
| Environment | LunarLander-v2 |
| Algorithm | PPO |
| Total timesteps | 100000 |
| Learning rate | 0.00025 |
| Number of environments | 8 |
| Steps per rollout | 128 |
| Gamma | 0.99 |
| GAE Lambda | 0.95 |
| Minibatches | 4 |
| Update epochs | 4 |
| PPO clip coefficient | 0.2 |
| Entropy coefficient | 0.01 |
| Value coefficient | 0.5 |
| Seed | 1 |
The trained agent was evaluated for 10 episodes.
The policy uses an Actor-Critic architecture.
The actor predicts action logits for the discrete LunarLander action space.
The critic estimates the value of the current state.
Both networks contain two fully connected hidden layers with 64 units and Tanh activation functions.
model.pt — trained PyTorch modelhyperparameters.txt — PPO hyperparametersevaluation.txt — evaluation resultsREADME.md — model card