Downloads · 30 days
0
vshwanilgv/wenavigatecontroller-ppo
wenavigatecontroller-ppo is a reinforcement learning model from vshwanilgv. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Low-level PPO controller for the WeNavigate Vision-Language Navigation system. Trained to execute VLM navigation commands (forward / turn-left / turn-right / stop) inside Facebook Habitat-sim with HM3D scenes.
Downloads · 30 days
0
Access
Public
Updated Apr 30, 2026
Repo size
70.5 MB
Likes
0
Public
Click a slice to open those files.
.pt35.3 MB · 99%
From the Hugging Face model README
Low-level PPO controller for the WeNavigate Vision-Language Navigation system. Trained to execute VLM navigation commands (forward / turn-left / turn-right / stop) inside Facebook Habitat-sim with HM3D scenes.
| Parameter | Value |
|---|---|
| n_rollout | 2048 |
| n_epochs | 4 |
| batch_size | 256 |
| lr | 0.0003 |
| gamma | 0.99 |
| gae_lambda | 0.95 |
| clip_eps | 0.2 |
| entropy_coef | 0.01 |
import torch
from ppo_policy import PPOPolicy
policy = PPOPolicy()
ckpt = torch.load("policy_update_XXXXX.pt", map_location="cpu")
policy.load_state_dict(ckpt["policy_state"])
policy.eval()
# obs: dict with keys depth (64,64), command (4,), proprioception (3,)
action, log_prob, entropy, value = policy.get_action_and_value(
depth.unsqueeze(0),
command.unsqueeze(0),
prop.unsqueeze(0),
)
Trained on wenavigatecontroller-long-episodes — HM3D minival scenes 00800–00809, 160 train episodes, 160 eval episodes.