Downloads · 30 days
73
81% of all-time downloads
DecisionFacts/Physical_AI_SO101_Cup_Nesting_ACT_Policy
Physical_AI_SO101_Cup_Nesting_ACT_Policy is a robotics model from DecisionFacts. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as cc-by-nc-4.0.
An Action Chunking Transformer (ACT) manipulation policy trained on real-world teleoperation data from the SO-101 robotic arm (sofollower) for a cup nesting task: picking up a cup and seating it inside a second cup.
Downloads · 30 days
73
81% of all-time downloads
All-time downloads
90
Public
Parameters
51.7M
207 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors207 MB · 100%
From the Hugging Face model README
An Action Chunking Transformer (ACT) manipulation policy trained on real-world teleoperation
data from the SO-101 robotic arm (so_follower) for a cup nesting task: picking up a cup
and seating it inside a second cup.
This is the policy counterpart to the DecisionFacts Physical AI SO-101 Cup Nesting dataset. It is released as a reference checkpoint — a demonstration of what a standard ACT recipe achieves on a single clean tranche of DecisionFacts teleoperation data, and a starting point for fine-tuning, benchmarking, and reproduction.
| Architecture | ACT (Action Chunking Transformer) — CVAE + transformer encoder/decoder over ResNet vision backbones |
| Framework | LeRobot |
| Parameters | ~51.7 M |
| Precision | FP32 (safetensors) |
| Embodiment | SO-101 follower arm, 6-DoF |
| Task | Cup nesting (grasp cup, align, nest into target cup) |
| Observation space | Dual RGB cameras (cam_front, cam_top) @ 480×640, 30 fps + 6-DoF joint state |
| Action space | 6-DoF joint targets (shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper) |
| Control rate | 30 Hz |
| License | CC BY-NC 4.0 (research and evaluation only) |
Why ACT. ACT predicts a chunk of future actions from a single observation rather than one step at a time. For contact-rich, precision-alignment tasks like nesting, this suppresses the compounding error and jittery re-planning that plague per-step behavioral cloning, and it produces smooth, committed trajectories at 30 Hz without a separate motion planner.
Trained on the DecisionFacts SO-101 Cup Nesting teleoperation tranche, recorded in LeRobotDataset v3.0 format.
| Stage | Episodes |
|---|---|
| Recorded | 200 |
| Excluded — camera-timestamp defects | 8 |
| Verified clean, used for training | 192 |
Every episode is an expert human teleoperation demonstration, captured with synchronized dual-camera RGB video and full proprioceptive state at 30 fps.
Quality control. Prior to training, all 200 episodes were screened for completeness, timing alignment, and coherent task execution. Eight episodes were removed for camera-timestamp defects — frame timestamps that drifted out of lock-step with the joint-state stream, which would inject misaligned vision/action pairs into training. The remaining 192 episodes were verified clean and used in full.
Domain randomization. Cup poses, positions, orientations, and workspace configuration were varied across episodes rather than repeating a fixed setup, so the policy sees a broader distribution of visual and spatial conditions than a single staged scene would provide.
Note on the public dataset repository. The linked dataset repo is a public evaluation sample of the DecisionFacts catalog and contains a subset of the episodes. This checkpoint was trained on the full 192-episode verified tranche. Exact reproduction of this checkpoint requires the full tranche — contact [email protected].
| Steps | 100,000 (full run, completed) |
| Hardware | 1 × NVIDIA L4 (Google Cloud Platform) |
| Wall-clock time | ~6 hours |
| Stability | Zero crashes on the cleaned dataset |
| Trainer | lerobot-train (LeRobot ACT recipe) |
Normalization statistics were taken from the dataset's meta/stats.json. The run completed end to
end with no restarts, NaN losses, or dataloader failures — the eight defective episodes were the
sole source of instability observed in earlier attempts, and removing them was sufficient to make
the run clean.
pip install lerobot
lerobot-train \
--policy.type=act \
--dataset.repo_id=DecisionFacts/Physical_AI_SO101_Cup_Nesting_Task \
--steps=100000 \
--batch_size=8 \
--output_dir=outputs/train/so101_cup_nesting_act \
--job_name=so101_cup_nesting_act \
--policy.device=cuda \
--wandb.enable=true
Hyperparameters not listed above follow LeRobot's default ACT configuration (chunk size,
optimizer, learning rate, backbone). Verify against config.json in this repository, which is
authoritative for this checkpoint.
from lerobot.policies.act.modeling_act import ACTPolicy
policy = ACTPolicy.from_pretrained(
"DecisionFacts/Physical_AI_SO101_Cup_Nesting_ACT_Policy"
)
policy.eval()
policy.to("cuda")
lerobot-record \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_so101 \
--robot.cameras="{ cam_front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, cam_top: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30} }" \
--dataset.repo_id=<your-hf-user>/eval_so101_cup_nesting \
--dataset.single_task="Nest the cup" \
--policy.path=DecisionFacts/Physical_AI_SO101_Cup_Nesting_ACT_Policy
Camera naming matters. The policy expects observation keys observation.images.cam_front and
observation.images.cam_top. If your camera keys differ, remap them before inference — silently
swapping the two viewpoints will degrade performance badly rather than fail loudly.
lerobot-train \
--policy.path=DecisionFacts/Physical_AI_SO101_Cup_Nesting_ACT_Policy \
--dataset.repo_id=<your-hf-user>/<your-dataset> \
--steps=20000 \
--output_dir=outputs/train/cup_nesting_finetune
Rollout success rates on physical hardware are not yet published for this checkpoint. We are running a standardized evaluation protocol (fixed trial count, randomized cup placements, held-out starting configurations) and will update this section with results, including per-stage breakdown (reach / grasp / align / nest) and failure-mode analysis.
Until then, treat this checkpoint as a reference training artifact, not a performance-validated product. If you evaluate it, we would like to hear what you find — open a discussion in the Community tab.
Intended for:
Not intended for:
Safety. This policy commands a real robot arm. Always maintain a clear workspace, enforce torque and joint limits, keep an emergency stop within reach, and supervise every rollout.
Released under CC BY-NC 4.0 — research and evaluation purposes only, consistent with the license of the underlying dataset.
Commercial use, redistribution, or deployment of this policy (or models derived from it) in commercial products requires a separate license. Because this checkpoint is derived from the DecisionFacts Physical AI dataset, dataset licensing terms flow through to the model weights.
This checkpoint is trained on a sample tranche of a larger, continuously growing dataset. The full catalog includes additional tasks, larger per-task episode counts, and can be scoped to specific manipulation skills, environments, or diversity requirements.
We offer:
To discuss access to the full catalog, custom collection, or commercial licensing, contact [email protected].
If you use this model in your work, please cite it as follows:
@misc{decisionfacts_so101_cup_nesting_act,
title = {DecisionFacts Physical AI Policy --- SO-101 Cup Nesting (ACT)},
credits = {Prabhu Raghav, Sreeram B Unni, Balamurugan Pandi, Sriram Gopalan},
year = {2026},
howpublished = {\url{https://huggingface.co/DecisionFacts/Physical_AI_SO101_Cup_Nesting_ACT_Policy}}
}
Please also cite the underlying dataset:
@misc{decisionfacts_teleops_dataset,
credits = {Prabhu Raghav, Sreeram B Unni, Balamurugan Pandi, Sriram Gopalan},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/DecisionFacts/Physical_AI_SO101_Cup_Nesting_Task}}
}
And the ACT method:
@inproceedings{zhao2023learning,
title = {Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware},
author = {Zhao, Tony Z. and Kumar, Vikash and Levine, Sergey and Finn, Chelsea},
booktitle = {Robotics: Science and Systems (RSS)},
year = {2023}
}
DecisionFacts Inc · Physical AI data and policies for real-world robot learning
</div>