Downloads · 30 days
10
10% of all-time downloads
tsilva/Level3-4_stable-baselines3-ppo_4141caab
Level3-4_stable-baselines3-ppo_4141caab is a reinforcement learning model from tsilva. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for stable-baselines3. The card lists the license as mit.
Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level3-4, trained and evaluated with rlab.
Downloads · 30 days
10
10% of all-time downloads
All-time downloads
99
Public
Repo size
21.4 MB
Likes
0
Public
Click a slice to open those files.
.zip21.1 MB · 98%
From the Hugging Face model README
Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level3-4, trained and evaluated with
rlab.
| Item | Value |
|---|---|
| Task | Complete SuperMarioBros-Nes-v0 Level3-4 |
| Provider | supermariobrosnes-turbo |
| Algorithm | ppo |
| Checkpoint | Step 3220272 |
| Evaluation | stochastic full evaluation, 100 episodes |
| Success | minimum 96.0%, mean 96.0% |
| Mean return | 2283.246 |
| Release | v1 |
| Preview | Root replay.mp4 |
| YouTube | Watch on YouTube |
git clone https://github.com/tsilva/rlab
cd rlab
git checkout a21e8fc154ecf3e47e39f1fc523398b575cd12ed
uv sync --frozen
Import the ROM, then play or evaluate the immutable checkpoint:
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level3-4_stable-baselines3-ppo_4141caab/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level3-4_stable-baselines3-ppo_4141caab/resolve/v1/model.zip
Action selection was stochastic under the published evaluation environment contract.
| Start | Episodes | Successes | Success rate | Mean return |
|---|---|---|---|---|
| Level3-4 | 100 | 96 | 96.0% | 2283.246 |
| Item | Value |
|---|---|
| Environment | supermariobrosnes-turbo:SuperMarioBros-Nes-v0 |
| Environment hash | sha256:7969db636994526e3c3c1614d378f6de882854a51d54ac9d3ac664fe00b4248a |
| Preprocessing | {"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0} |
| Action contract | {"set":"simple"} |
| Item | Value |
|---|---|
| Source | rlab |
| Run | Level3-4_base_s1_20260704T131342Z |
| Recipe | base |
| Seed | 1 |
| Source commit | a21e8fc154ecf3e47e39f1fc523398b575cd12ed |
| Evaluated artifact | tsilva/SuperMarioBros-Nes-v0/Level3-4_base_s1_20260704T131342Z-final:latest |
| File | Purpose |
|---|---|
model.zip | Stable-Baselines3 policy checkpoint |
model.json | Versioned checkpoint identity, policy type, provenance, and recipe binding |
recipe.json | Versioned execution and evaluation contract |
release_manifest.json | Release identity, evaluation evidence, and artifact hashes |
replay.mp4 | Browser-safe representative episode |
LICENSE | License for rlab-authored policy weights and publication material |
Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.
The rlab-authored policy weights and publication material are licensed under the MIT License in
LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.
This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.
Stable-Baselines3PPOstable_baselines3.ppo.ppo.PPO4141caabd238d66ed39cf0949c208ba5661d8fc16cc20eb1928ee1d24f8fe369hf://tsilva/Level3-4_stable-baselines3-ppo_4141caab@v1checkpoint-3220272The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.