Downloads · 30 days
0
zuzuzzy/NS-VLA
NS-VLA is a robotics model from zuzuzzy. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Mar 10, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md2.4 KB · 62%
From the Hugging Face model README
NS-VLA is a neuro-symbolic Vision-Language-Action framework that combines symbolic reasoning with neural control for robotic manipulation. The model introduces:
| Property | Value |
|---|---|
| Architecture | Qwen3-VL-2B + Symbolic Classifier + Action Generator |
| Parameters | ~2B (VLM backbone frozen) |
| Training | Stage I: Supervised Pretraining → Stage II: Online RL (GRPO) |
| Input | RGB image (224×224) + natural language instruction |
| Output | Continuous end-effector actions (chunked, H=8) |
| Benchmark | Setting | Success Rate (%) |
|---|---|---|
| LIBERO | Full demonstrations | 98.6 |
| LIBERO | 1-shot (one demo per task) | 69.1 |
| LIBERO-Plus | Zero-shot generalization | 79.4 |
| CALVIN ABC→D | Zero-shot 5-task chain | 91.2 |
⚠️ Note: Model weights will be released upon paper acceptance. Please check back soon.
# Example usage (coming soon)
from nsvla import NSVLAAgent
agent = NSVLAAgent.from_pretrained("Zuzuzzy/NS-VLA")
action = agent.predict(image=obs, instruction="pick up the red mug")
@article{zhu2026nsvla,
title={NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models},
author={Zhu, Ziyue and Wu, Shangyang and Zhao, Shuai and Zhao, Zhiqiu and Li, Shengjie and Wang, Yi and Li, Fang and Luo, Haoran},
journal={arXiv preprint arXiv:XXXX.XXXXX},
year={2026}
}
This model is released under the Apache 2.0 License.