Downloads · 30 days
0
amaldo123/ppo-spaceinvaders-scratch
ppo-spaceinvaders-scratch is a reinforcement learning model from amaldo123. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product.
This model was trained as part of a project building policy-gradient methods up from first principles: REINFORCE - A2C - A2C+GAE - PPO (this checkpoint).
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
6.8 MB
Likes
0
Public
Click a slice to open those files.
.pt6.8 MB · 100%
From the Hugging Face model README
This model was trained as part of a project building policy-gradient methods up from first principles: REINFORCE -> A2C -> A2C+GAE -> PPO (this checkpoint).
ALE/SpaceInvaders-v5, preprocessed with grayscale + 84x84 resize + 4-frame skip +
4-frame stack + terminal-on-life-loss (standard DQN-style preprocessing).Nature CNN trunk (3 conv layers -> 512-d FC) with two heads: policy logits (6 actions) and a scalar state-value.
Mean return over 10 greedy evaluation episodes: 105.0 (std 0.0)
import torch
from model_def import ActorCritic # see training notebook for the class definition
net = ActorCritic(n_actions=6)
net.load_state_dict(torch.load("ppo_spaceinvaders.pt"))
net.eval()