Downloads · 30 days
0
Aathi07/ppo-LunarLander-v2
ppo-LunarLander-v2 is a reinforcement learning model from Aathi07. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product.
This is a from-scratch PyTorch implementation of PPO trained for Unit 8 Part 1 of the Hugging Face Deep RL Course.
Downloads · 30 days
0
Access
Public
Updated Sep 12, 2026
Repo size
633 KB
Likes
0
Public
Click a slice to open those files.
.mp4198 KB · 81%
From the Hugging Face model README
This is a from-scratch PyTorch implementation of PPO trained for Unit 8 Part 1 of the Hugging Face Deep RL Course.
Mean reward: 30.00 +/- 52.96 over 10 evaluation episodes.
Load model.pt into the Agent class defined in the training script and call
get_action_and_value on an observation tensor.