Downloads · 30 days
0
AMD-PAVS-AI/smolVLA
smolVLA is a robotics model from AMD-PAVS-AI. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
754 KB
Likes
1
Public
Click a slice to open those files.
.png754 KB · 99%
From the Hugging Face model README

SmolVLA (vision-language-action) is a behavior-cloning policy from Hugging Face LeRobot for 6-DOF robot arm control. This repository packages inference for robot arm action prediction using PyTorch, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of SmolVLA found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the smolVLA AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task: Robot arm action prediction (vision-language-action)
Dataset: BlankHead/so101_redcube_greencloth_3cams (LeRobot format)
Output metrics: MAE, RMSE (per-joint and per-episode)
PyTorch note: CPU runs FP32; GPU runs BF16. No NPU (VitisAI) path is available —
make setup-npu,make benchmark-npu, andmake evaluate-npuprint an informational note and exit cleanly.
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| PyTorch | FP32 | HIP (ROCm) | AMD CPU | — |
| PyTorch | BF16 | HIP (ROCm) | AMD Instinct™ / Radeon™ GPU | No NPU path available |
For setup instructions, evaluation scripts, and custom configuration options, see the smolVLA on GitHub.
Model Type: Vision-language-action policy for robot arm control
Base Model: lerobot/smolvla_base
Model Stats:
rtc_config.enabled (Real-Time Chunking), num_steps (flow-matching denoising passes per chunk), n_action_steps (actions consumed per chunk)Open-loop offline evaluation is fully implemented: make evaluate-<device> runs inference on recorded dataset episodes and computes per-joint and per-episode MAE / RMSE against the recorded ground-truth actions. Lower is better for both metrics.
| Metric | Description |
|---|---|
| MAE | Mean Absolute Error — average absolute difference between predicted and ground-truth joint positions across all timesteps. Lower is better. |
| RMSE | Root Mean Squared Error — penalizes large deviations more heavily than MAE. Lower is better. |
Published Results — Dataset: BlankHead/so101_redcube_greencloth_3cams (13 episodes, chunked_rtc mode):
| Metric | Value |
|---|---|
| Average MAE | 3.9521 |
| Average RMSE | 7.7251 |
Per-joint breakdown:
| Joint | Avg MAE | Avg RMSE |
|---|---|---|
| shoulder_pan | 3.3654 | 5.1317 |
| shoulder_lift | 7.8224 | 14.0630 |
| elbow_flex | 4.9989 | 8.7409 |
| wrist_flex | 2.4138 | 3.4653 |
| wrist_roll | 2.6728 | 3.9475 |
| gripper | 2.4394 | 4.5340 |
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes: