Downloads ยท 30 days
0
Punit71/firefighter-gridworld-leaderboard
firefighter-gridworld-leaderboard is a reinforcement learning model from Punit71. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for custom. The card lists the license as mit.
A reinforcement learning benchmark in a 4x4 grid world where the agent must:
Downloads ยท 30 days
0
Access
Public
Updated May 5, 2025
Repo size
421 KB
Likes
0
Public
Click a slice to open those files.
.gif421 KB ยท 85%
From the Hugging Face model README
A reinforcement learning benchmark in a 4x4 grid world where the agent must:
The environment features deterministic and stochastic versions with discrete actions, rewards, penalties, and sprite-based rendering.
| Rank | Model | Mean Reward | Std Dev | Success Rate | Notes |
|---|---|---|---|---|---|
| ๐ฅ 1 | MCTS | 27.4 | 3.64 | 1.00 | 50 simulations, random rollout |
| 2 | PPO | 4.0 | 5.83 | ~0.40 | Trained with Stable-Baselines3 |
| 3 | DQN | โ30.0 | 23.9 | โ | Failed task consistently |
Each agent is evaluated over 300 episodes
Maximum steps per episode: 60
Environment starts with the robot in the top-left
Rewards:
pip install -r requirements.txt
git clone https://huggingface.co/spaces/YOUR_USERNAME/firefighter-gridworld-leaderboard
cd firefighter-gridworld-leaderboard
python evaluation/evaluate_custom_agent.py --path ./my_agent.zip --algo PPO
eval_results.json via Pull Request.Custom environment follows Gymnasium standards:
import gymnasium as gym
from env.firefighter_env import FireFighterEnv
env = FireFighterEnv()
obs, info = env.reset()
for _ in range(60):
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
break
Include in your Pull Request:
eval_results.jsonenv/ โ environment codeagents/ โ training scripts (PPO, DQN, MCTS)evaluation/ โ evaluation and renderingmodels/ โ saved agentsassets/ โ sprites and animationMIT License. Contributions welcome!