Downloads · 30 days
5
7% of all-time downloads
tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc
Level4-1_stable-baselines3-ppo_fe8a02dc is a reinforcement learning model from tsilva. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for stable-baselines3. The card lists the license as mit.
Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level4-1, trained and evaluated with rlab.
Downloads · 30 days
5
7% of all-time downloads
All-time downloads
76
Public
Repo size
21.9 MB
Likes
0
Public
Click a slice to open those files.
.zip21.1 MB · 96%
From the Hugging Face model README
Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level4-1, trained and evaluated with
rlab.
| Item | Value |
|---|---|
| Task | Complete SuperMarioBros-Nes-v0 Level4-1 |
| Provider | supermariobrosnes-turbo |
| Algorithm | ppo |
| Checkpoint | Step 7500000 |
| Evaluation | stochastic full evaluation, 100 episodes |
| Success | minimum 97.0%, mean 97.0% |
| Mean return | 3529.275 |
| Release | v1 |
| Preview | Root replay.mp4 |
| YouTube | Watch on YouTube |
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 044dcc3d2a978fb1a77524822238c4c7814c002d
uv sync --frozen
Import the ROM, then play or evaluate the immutable checkpoint:
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
Action selection was stochastic under the published evaluation environment contract.
| Start | Episodes | Successes | Success rate | Mean return |
|---|---|---|---|---|
| Level4-1 | 100 | 97 | 97.0% | 3529.275 |
| Item | Value |
|---|---|
| Environment | supermariobrosnes-turbo:SuperMarioBros-Nes-v0 |
| Environment hash | sha256:fa0f3dd473fac528938a2cc2130dfe2d516ef164946b6545f2262ff6777c4bef |
| Preprocessing | {"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0} |
| Action contract | {"set":"simple"} |
| Item | Value |
|---|---|
| Source | rlab |
| Run | Level4-1_base_s1_20260704T182045Z |
| Recipe | base |
| Seed | 1 |
| Source commit | 044dcc3d2a978fb1a77524822238c4c7814c002d |
| Evaluated artifact | tsilva/SuperMarioBros-Nes-v0/Level4-1_base_s1_20260704T182045Z-checkpoint:step-7500000 |
| File | Purpose |
|---|---|
model.zip | Stable-Baselines3 policy checkpoint |
model.json | Versioned checkpoint identity, policy type, provenance, and recipe binding |
recipe.json | Versioned execution and evaluation contract |
release_manifest.json | Release identity, evaluation evidence, and artifact hashes |
replay.mp4 | Browser-safe representative episode |
LICENSE | License for rlab-authored policy weights and publication material |
Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.
The rlab-authored policy weights and publication material are licensed under the MIT License in
LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.
This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.
Stable-Baselines3PPOstable_baselines3.ppo.ppo.PPOfe8a02dca01803a336dfecd8803be53541b192c0b0d05a525b4a6afcbae4c58ahf://tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc@v1checkpoint-7500000The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.