Downloads · 30 days
0
abdul004/pi05_so101_checkpoint
pi05_so101_checkpoint is a robotics model from abdul004. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as apache-2.0.
A fine-tuned Pi0.5 (π₀.₅) Vision-Language-Action model for the ball-in-cup task using the SO-101 robot arm.
Downloads · 30 days
0
Access
Public
Updated Feb 2, 2026
Repo size
134 GB
Likes
0
Public
Click a slice to open those files.
Other89.4 GB · 100%
From the Hugging Face model README
A fine-tuned Pi0.5 (π₀.₅) Vision-Language-Action model for the ball-in-cup task using the SO-101 robot arm.
Goal: Pick up an orange ball from the table and place it into a pink cup.
Robot: SO-101 - 6-DOF robot arm with gripper
Cameras: Dual camera setup (overhead + wrist-mounted)
Pi0.5 is a Vision-Language-Action (VLA) model from Physical Intelligence:
| Component | Description |
|---|---|
| Vision Encoder | SigLIP 400M - processes camera images |
| Language Model | Gemma 2B - scene understanding & task grounding |
| Action Expert | Flow Matching head - generates smooth action trajectories |
| Total Parameters | ~3B |
The model takes natural language instructions + camera images → outputs continuous joint actions.
| Parameter | Value |
|---|---|
| Base Model | Pi0.5 (Physical Intelligence) |
| Dataset | abdul004/so101_ball_in_cup_v5 |
| Episodes | 72 teleoperated demonstrations |
| Frames | 25,045 |
| Fine-tuning Steps | 5,000 |
| Hardware | A100 80GB on RunPod |
| Training Time | ~3-4 hours |
| Cost | ~$6-8 USD |
| Framework | OpenPi (JAX/Flax) |
We implemented JPEG compression to reduce network transfer time for remote inference:
| Location | Raw Images | JPEG (Q80) | Speedup |
|---|---|---|---|
| EU Spot | 1448ms | 375ms | 3.9x |
| US On-Demand | 600ms | 270ms | 2.2x |
| Metric | Before | After |
|---|---|---|
| Payload Size | 1.8 MB | 71 KB |
| Control Rate (US) | 1.7 Hz | 3.7 Hz |
| Compression Ratio | - | 25x |
[RunPod GPU Server] [Robot Mac]
┌─────────────────┐ ┌──────────────┐
│ Pi0.5 Model │◄── WSS ────►│ run_pi05.py │
│ (RTX 4090) │ JPEG │ (Robot ctrl) │
└─────────────────┘ └──────────────┘
Side-by-side: Overhead camera (left) + Wrist camera (right) - Smooth 3.7 Hz control
Side-by-side: Same task but with raw image transfer - 1.7 Hz control
5-frame composite: Start → Approach → Grasp → Transport → Final
Same task without JPEG optimization
# Clone OpenPi fork with JPEG support
git clone https://github.com/abdulrahman004/openpi.git
cd openpi
uv sync
# Download checkpoint
uv run huggingface-cli download abdul004/pi05_so101_checkpoint \
--include "4999/**" \
--local-dir checkpoints/pi05_so101
# Start server
uv run scripts/serve_policy.py --port 8000 \
policy:checkpoint \
--policy.config=pi05_so101 \
--policy.dir=checkpoints/pi05_so101/4999
pip install openpi-client
# Run inference with JPEG compression
python run_pi05.py --server wss://YOUR-POD-8000.proxy.runpod.net
# Or without compression (slower)
python run_pi05.py --server wss://YOUR-POD-8000.proxy.runpod.net --no-jpeg
Real-world demonstrations recorded externally during evaluation runs:
External phone recording showing smooth robot control with JPEG compression
Same task without compression - noticeably slower/choppier control
Ball placed at workspace edge - a position that appeared in <10% of training episodes
Trained on the same dataset:
| Policy | Architecture | Inference | Grasp | Generalization |
|---|---|---|---|---|
| Pi0.5 | VLA (3B params) | Remote GPU | ✅ | ✅ Edge positions |
| ACT | Transformer (25M) | Local | ✅ | ❌ Center only |
ACT failed at edge positions - the policy was only trained with ~72 episodes where the ball was mostly in the center/reachable area. When the ball was placed at the edge of the workspace, ACT would miss or fail to reach it entirely.
Pi0.5 succeeds at edge positions despite having the same training data. This demonstrates the power of VLA pre-training:
The base Pi0.5 model was trained on data from many different robot arms performing various tasks. This gives it a strong prior on reachable workspace and arm kinematics that ACT (trained from scratch) simply doesn't have.
Remote Inference Setup:
Known Issues:
@misc{so101_pi05_ball_in_cup,
author = {Abdul},
title = {SO-101 Ball-in-Cup Pi0.5 Fine-tuning},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/abdul004/pi05_so101_checkpoint}
}