Downloads · 30 days
9
17% of all-time downloads
io-intelligence/smolvla_so101_stack_cups
smolvla_so101_stack_cups is a robotics model from io-intelligence. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as apache-2.0.
SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.
Downloads · 30 days
9
17% of all-time downloads
All-time downloads
54
Public
Parameters
450M
907 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors907 MB · 100%
How the weights are stored.
BF16447M · 99%
From the Hugging Face model README
SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.
Fine-tuned on SO-101 for the task "Stack the cups".
<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/640e21ef3c82bd463ee5a76d/aooU0a3DMtYmy_1IWMaIM.png" alt="smolvla architecture" width="85%"/> </p> <p align="center"> <video src="https://huggingface.co/io-intelligence/smolvla_so101_stack_cups/resolve/main/inference_demo.mp4" controls width="70%"></video> </p> <p align="center"><em>Real-robot inference demo (also available as <code>inference_demo.mp4</code> in this repo).</em></p>This policy has been trained and pushed to the Hub using LeRobot.
Learn how to train and run it in the LeRobot smolvla guide, or browse the full documentation.
so101_follower (SO-101)The dataset records views as front / top / wrist. During SmolVLA training they are renamed to camera1 / camera2 / camera3. At inference you should keep the physical names on the robot and pass the same rename_map.
| Physical camera | Mount / role | Policy feature after rename |
|---|---|---|
front | Front-facing view of the workspace (Astra) | observation.images.camera1 |
top | Top-down / overhead view (Astra) | observation.images.camera2 |
wrist | Wrist / gripper camera | observation.images.camera3 |
{
"observation.images.front": "observation.images.camera1",
"observation.images.top": "observation.images.camera2",
"observation.images.wrist": "observation.images.camera3"
}
The policy consumes these observation features and produces these action features.
Inputs
| Feature | Physical source | Type | Shape |
|---|---|---|---|
observation.state | Joint state | STATE | (6,) |
observation.images.camera1 | front | VISUAL | (3, 256, 256) |
observation.images.camera2 | top | VISUAL | (3, 256, 256) |
observation.images.camera3 | wrist | VISUAL | (3, 256, 256) |
Outputs
| Feature | Type | Shape |
|---|---|---|
action | ACTION | (6,) |
| Setting | Value |
|---|---|
| Training steps | 30000 |
| Batch size | 8 |
| Optimizer | adamw |
| Learning rate | 0.0001 |
| Seed | 1000 |
| LeRobot version | 0.6.1 |
New to LeRobot? These guides cover the full workflow:
lerobot package.lerobot-* commands.Use physical camera keys front / top / wrist, then apply the rename map so they match the policy's camera1 / camera2 / camera3 features. SmolVLA works best with RTC inference.
lerobot-rollout \
--strategy.type=base \
--robot.type=so101_follower \
--robot.port=<your_robot_port> \
--robot.cameras="{ \
front: {type: opencv, index_or_path: <front_device>, width: 640, height: 480, fps: 30}, \
top: {type: opencv, index_or_path: <top_device>, width: 640, height: 480, fps: 30}, \
wrist: {type: opencv, index_or_path: <wrist_device>, width: 640, height: 480, fps: 30} \
}" \
--policy.path=io-intelligence/smolvla_so101_stack_cups \
--rename_map='{"observation.images.front":"observation.images.camera1","observation.images.top":"observation.images.camera2","observation.images.wrist":"observation.images.camera3"}' \
--inference.type=rtc \
--task="Stack the cups" \
--duration=60
Replace <your_robot_port> and the three camera device paths with your machine values. Camera names must stay front / top / wrist (not camera1/2/3 on the robot side).
When --strategy.type=base is used the script doesn't record episodes. Set --duration=0 (or omit duration depending on your CLI) to run until Ctrl+C. For more information see the rollout / inference docs.
This policy type is usually fine-tuned from the pretrained base model lerobot/smolvla_base:
lerobot-train \
--dataset.repo_id=${HF_USER}/<dataset> \
--policy.path=lerobot/smolvla_base \
--output_dir=outputs/train/<policy_repo_id> \
--job_name=lerobot_training \
--policy.device=cuda \
--policy.repo_id=${HF_USER}/<policy_repo_id> \
--wandb.enable=true
Writes checkpoints to outputs/train/<policy_repo_id>/checkpoints/.
Real-robot inference demo of stacking cups is included as inference_demo.mp4 (copied from the training dataset root).
| Task | Notes |
|---|---|
| Stack the cups | Qualitative success demo on SO-101 (see video above) |
If you use this policy, please cite the method linked in the description above, along with LeRobot:
@misc{cadene2024lerobot,
author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
howpublished = "\url{https://github.com/huggingface/lerobot}",
year = {2024}
}