Downloads · 30 days
5
4% of all-time downloads
aryankakad/cupstack_bspline_act
cupstack_bspline_act is a robotics model from aryankakad. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as apache-2.0.
An ACT policy that predicts B-spline trajectory segments instead of discrete action chunks, trained on SO-101 cup stacking. This is the "Reg.+BSP" variant from B-spline Policy: Accelerating Manipulation Policies via B…
Downloads · 30 days
5
4% of all-time downloads
All-time downloads
130
Public
Parameters
51.6M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 100%
From the Hugging Face model README
An ACT policy that predicts B-spline trajectory segments instead of discrete action chunks, trained on SO-101 cup stacking. This is the "Reg.+BSP" variant from B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations.
The practical consequence: you can run this checkpoint faster without retraining it.
The policy outputs a curve, so executing a(n·t) replays the same trajectory geometry
n times faster. Speed-up is an inference flag.
[!WARNING] This checkpoint has never been evaluated on held-out data or on hardware. It was trained on all 50 episodes with no validation split, so its numbers measure fit, not generalization. Treat first hardware runs as untested — keep the e-stop within reach. See Limitations.
lerobot-rollout \
--strategy.type=base \
--policy.path=aryankakad/cupstack_bspline_act \
--policy.speed_up=1.0 \
--device=cuda \
--robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=FOLLOWER \
--robot.cameras='{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, gripper: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}}' \
--task="stack the cups" \
--fps=30 --duration=30 --display_data=true
Raise --policy.speed_up to 2.0 or 4.0 to execute faster. Start at 1.0 and step up
only once the task succeeds.
The policy requires exactly these observation keys, so the --robot.cameras dict keys
must be front and gripper, lowercase:
| Key | Shape |
|---|---|
observation.images.front | (3, 480, 640) |
observation.images.gripper | (3, 480, 640) |
observation.state | (6,) |
Joint order is shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper.
Swapping the two cameras produces confident but wrong behaviour rather than an error.
This policy type is not in upstream LeRobot. You need a checkout containing the
bspline_act policy, installed with the scipy extra:
uv pip install -e ".[feetech,bspline]"
scipy>=1.15 is required for interpolate.generate_knots.
Unnormalization statistics are stored as buffers inside model.safetensors, so the
checkpoint is self-contained. You do not need the training dataset on the robot machine.
| Steps | 100,000 |
| Batch size | 32 |
| Epochs | 133.09 |
| Dataset | aryankakad/CUPSTACKING → B-spline converted |
| Episodes / frames | 50 / 24,044 @ 30 fps |
| Optimizer | AdamW, lr 2e-5 (backbone 1e-5), wd 1e-4 |
| Precision | bf16 autocast + channels_last |
| Hardware | 1× RTX 4090, 5 h 24 min |
| Seed | 1000 |
Learning rate is 2e-5 rather than ACT's default 1e-5, sqrt-scaled for batch 32.
| Step | loss | l1_loss | kld_loss |
|---|---|---|---|
| 500 | 3.027 | 0.390 | 0.264 |
| 10k | 0.166 | 0.104 | 0.007 |
| 50k | ~0.111 | 0.041 | 0.007 |
| 100k | 0.095 | 0.026 | 0.007 |
Training loss only — there was no validation split.
Standard ACT with one change: the decoder emits a B-spline parameter matrix rather than an action chunk.
(n_knots, 1 + action_dim) = (16, 7) = 112 values per prediction
column 0 knot vector, in source-frame units, 0 = "now"
columns 1: control points, one per joint
| Params | 52 M |
| Vision backbone | ResNet18 (ImageNet init) |
| dim_model / chunk_size | 512 / 16 |
| VAE | enabled, kl_weight 10.0 |
| B-spline degree | 3 (cubic, C² continuous) |
| bspline_chunk_size | 10 |
| Fitting tolerance ε | 0.2 (degrees) |
Knot spacing is fitted adaptively per episode, so a fixed 16 rows covers a variable time horizon — 0.53 s to 1.47 s per segment on this dataset, 0.84 s mean. The network only runs when a segment is exhausted, which decouples policy rate from control rate.
ε = 0.2 is not the paper's value. The paper uses 0.002 for metre-scale end-effector actions; SO-101 stores joint targets in degrees (~±120), roughly 100× larger. At ε = 0.002 compression is 1.06× — one knot per frame, which defeats the representation.
| ε | Compression | p99 reconstruction | Segment span |
|---|---|---|---|
| 0.05 | 1.4× | 0.05° | 0.44 s |
| 0.2 | 2.7× | 0.19° | 0.84 s |
| 0.5 | 4.1× | 0.48° | 1.24 s |
2.7× sits inside the paper's reported 1.12×–3.34× range.
Offline only. No hardware evaluation has been performed.
Open-loop prediction error — decoded trajectory vs ground-truth actions, measured on training data:
| Checkpoint | Mean | Median | p90 | Max |
|---|---|---|---|---|
| 20k | 5.99° | 4.12° | 8.78° | 45.2° |
| 50k | 3.34° | 2.28° | 6.53° | 30.8° |
| 100k | 2.10° | 1.61° | 3.86° | 18.6° |
Error was still falling at 100k and 2.10° is far from zero, which argues against outright memorization — but this is training data, so it is not evidence of generalization.
The representation is not the bottleneck: B-spline fitting reconstructs to 0.19° p99, so essentially all of the 2.10° is policy prediction error.
Temporal rescaling is exact. On a real predicted segment (0.76 s span), a(2t) and
a(4t) reproduce the 1× samples to 0.00e+00, consuming the segment in 22 / 11 / 5
control ticks.
eval_steps: 0).
Generalization is unmeasured.@article{han2026b,
title={B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations},
author={Han, Xiaoshen and Xiong, Haoyu and Chen, Haonan and Liu, Chaoqi and
Torralba, Antonio and Zhu, Yuke and Du, Yilun},
journal={arXiv preprint arXiv:2607.09648},
year={2026}
}