Downloads · 30 days
0
physicalai-bmi/nano-vla-stack
nano-vla-stack is a robotics model from physicalai-bmi. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for safetensors. The card lists the license as cc-by-4.0.
A multi-step manipulation vision-language-action policy that runs fully in your browser. Instruction: "stack the red on the green." The arm grasps the named source block, carries it, and places it on the target block…
Downloads · 30 days
0
Access
Public
Updated Jul 12, 2026
Parameters
55.7K
223 KB on disk
Likes
0
Public
Click a slice to open those files.
.json1.2 MB · 84%
From the Hugging Face model README
A multi-step manipulation vision-language-action policy that runs fully in your browser. Instruction: "stack the red on the green." The arm grasps the named source block, carries it, and places it on the target block — a real two-phase task, not a single reach. Run it live at https://physicalai-bmi.org/research/vla.
| metric | value |
|---|---|
| Stack success (grasp → carry → place the correct block) | 94.0% |
| Grasp rate | 94.2% |
| Success with the source instruction flipped | 22.8% |
The flipped row shows the language matters: tell it to move the other block and it does, so success measured against the original target drops sharply. Trained by behavior cloning on a Jacobian-IK two-phase expert (100% ceiling) with DART state-noise injection to survive closed-loop drift across both phases.
Files: model.safetensors, vla.web.json (float32 for in-browser; forward verified
7.6e-8 vs safetensors), metrics.json. CC-BY-4.0, Institute for Physical AI @ BMI.
See also nano-vla-arm-3d (perspective reach) and nano-vla-reach.