Downloads · 30 days
0
maxmill/aq-mario
aq-mario is a reinforcement learning model from maxmill. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
This repository contains a trained action-conditioned latent world model from the full3eppureema999 experiment in AQ-Mario. It uses an EMA-JEPA architecture to learn the dynamics of Super Mario Bros. 1-1 from frame se…
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2026
Repo size
141 MB
Likes
1
Trending 1
Click a slice to open those files.
.pt140 MB · 98%
From the Hugging Face model README
This repository contains a trained action-conditioned latent world model from
the full3ep_pure_ema999 experiment in
AQ-Mario. It uses an EMA-JEPA
architecture to learn the dynamics of Super Mario Bros. 1-1 from frame
sequences and actions. The model predicts future latent states rather than RGB
frames, making it suitable for representation analysis and model-based planning
experiments.
This is a rollout preview from the EMA-JEPA model:
<video controls muted autoplay loop src="https://huggingface.co/maxmill/aq-mario/resolve/main/pure_jepa_ema.mp4"></video>
The complete training dataset is available at maxmill/aq-mario-smb1.
jepa.pt: PyTorch model checkpoint containing the JEPA world model, auxiliary
heads, EMA teacher, and AdamW optimizer state.metrics.jsonl: training metrics.gate_by_epoch.json: per-checkpoint probe and health measurements.gates.json: final representation and action-conditioning gate results.param_count.json: parameter report.ema_target=0.999)The final gate report is included for transparency. This run is not presented as a solved Mario controller: its final x/y/scroll probes and action gate do not pass the project's thresholds. See the GitHub repository for the loader, training code, evaluation protocol, and dataset documentation.
import torch
checkpoint = torch.load("jepa.pt", map_location="cpu", weights_only=False)
state_dict = checkpoint["jepa"]
The repository's aqmario.model.load_jepa helper can load the checkpoint when
the AQ-Mario source tree and configuration are available.
The training data consists of derived gameplay observations from Super Mario Bros. 1-1. AQ-Mario is an independent research project and is not affiliated with or endorsed by Nintendo. The source code is MIT licensed; this checkpoint and the derived gameplay data should be used subject to the rights and terms applicable to the underlying game content.