Downloads · 30 days
18
100% of all-time downloads
wangzhen-w/PanoVLN_base
PanoVLN_base is a image-text-to-text model from wangzhen-w. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
PanoVLN is a vision-and-language navigation policy that follows instructions from 360° RGB observations. It combines a Qwen3.5-4B backbone with PanoVGGT geometry features and predicts 18-action sequences. Confidence-g…
Downloads · 30 days
18
100% of all-time downloads
All-time downloads
18
Public
Parameters
6B
12 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors12 GB · 100%
From the Hugging Face model README
PanoVLN is a vision-and-language navigation policy that follows instructions from 360° RGB observations. It combines a Qwen3.5-4B backbone with PanoVGGT geometry features and predicts 18-action sequences. Confidence-guided execution selects how far to move before observing and planning again.
This checkpoint is intended for simulation and real-world navigation experiments. For more details on training, evaluation, and deployment, please refer to the repository documentation.