Downloads · 30 days
0
AMD-PAVS-AI/pi0.5
pi0.5 is a robotics model from AMD-PAVS-AI. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as gemma.
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
119 KB
Likes
0
Public
Click a slice to open those files.
.png119 KB · 95%
From the Hugging Face model README

pi0.5 (PI05Policy, Physical Intelligence) is a vision-language-action policy from Hugging Face LeRobot for 6-DOF robot arm control. This repository packages inference for robot arm action prediction using PyTorch, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs.
This is based on the implementation of pi0.5 found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the pi0.5 AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task: Robot arm action prediction (vision-language-action)
Dataset: BlankHead/so101_redcube_greencloth_3cams (LeRobot format)
Output metrics: MAE, RMSE (per-joint and per-episode)
Model variants: No
MODEL_SIZEvariants — the pipeline auto-detects PEFT/LoRA checkpoints viaadapter_config.jsonand merges the adapter, patchingaction_dimto the checkpoint's value.
PyTorch note: GPU only (ROCm — BF16); weights are streamed directly from disk into GPU VRAM in bf16, bypassing any CPU copy. A GPU with ≥32 GB VRAM is required — no CPU or NPU (VitisAI) path is available, since loading the ~14 GB fp32 weights on CPU would require >48 GB of system RAM. Requires Hugging Face auth (
HF_TOKEN) since the tokenizer pulls the gatedgoogle/paligemma-3b-pt-224repo.
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm. Key points:
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| PyTorch | BF16 | HIP (ROCm) | AMD Instinct™ / Radeon™ GPU (≥32 GB VRAM) | Direct-to-GPU weight streaming; no CPU/NPU path |
For setup instructions, evaluation scripts, and custom configuration options, see the pi0.5 on GitHub.
Model Type: Vision-language-action policy for robot arm control
Base Model: lerobot/pi05_base (PI05Policy)
Model Stats:
rtc_config.enabled (Real-Time Chunking), num_steps (flow-matching denoising passes per chunk), n_action_steps (actions consumed per chunk)Open-loop offline evaluation is fully implemented: make evaluate-gpu runs inference on recorded dataset episodes and computes per-joint and per-episode MAE / RMSE against the recorded ground-truth actions. Lower is better for both metrics.
| Metric | Description |
|---|---|
| MAE | Mean Absolute Error — average absolute difference between predicted and ground-truth joint positions across all timesteps. Lower is better. |
| RMSE | Root Mean Squared Error — penalizes large deviations more heavily than MAE. Lower is better. |
Published Results — filled from runs/eval/<dataset_tag>/loss.json:
| Dataset | Mode | Average MAE | Average RMSE |
|---|---|---|---|
| BlankHead/so101_redcube_greencloth_3cams (1 episode) | select_action | 19.5420 | 25.2683 |
| HarikrishnaVydana/Lerobot-redcube-wite-bg-pickdrop-v3 (1 episode) | chunked_rtc | 19.2629 | 32.3233 |
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes: