Downloads · 30 days
19
100% of all-time downloads
Shirley6/lunarlander
lunarlander is a reinforcement learning model from Shirley6. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for stable-baselines3.
使用 Stable-Baselines3 在 Gymnasium 的 LunarLander-v3 环境上训练的一个月球着陆器智能体(PPO 算法)。
Downloads · 30 days
19
100% of all-time downloads
All-time downloads
19
Public
Repo size
146 KB
Likes
0
Public
Click a slice to open those files.
.zip146 KB · 80%
From the Hugging Face model README
使用 Stable-Baselines3 在 Gymnasium 的 LunarLander-v3 环境上训练的一个月球着陆器智能体(PPO 算法)。
LunarLander-v3(Gymnasium / Box2D)| 超参数 | 值 |
|---|---|
| 算法 | PPO(MlpPolicy) |
| gamma | 0.999 |
| 总训练步数 | 1,500,000 |
| 并行环境数 | 16 |
| n_steps | 1024 |
| 学习率 | 3e-4 |