Downloads · 30 days
15
34% of all-time downloads
ted88168/colorlogo_value_round1
colorlogo_value_round1 is a robotics model from ted88168. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as apache-2.0.
This is the trajectory-value model used in the first Pi0.5 Advantage-Conditioned Policy (ACP) training round for an SO-101 color-block manipulation task.
Downloads · 30 days
15
34% of all-time downloads
All-time downloads
44
Public
Parameters
1.1B
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
How the weights are stored.
BF161.1B · 100%
From the Hugging Face model README
This is the trajectory-value model used in the first Pi0.5 Advantage-Conditioned Policy (ACP) training round for an SO-101 color-block manipulation task.
It is not a robot-control policy and should not be passed directly to lerobot-rollout. It estimates task-progress value for offline value, n-step advantage, and binary ACP-label inference.
The experimental ACP loop is:
The implementation and command templates are open source at BurningDawn8888/lerobot-pi05-acp. This is an experimental extension built on Hugging Face LeRobot and is not an official Pi0.5 feature.
| Setting | Value |
|---|---|
| Training steps | 8,000 |
| Batch size | 8 |
| Optimizer | AdamW |
| Learning rate | 5e-5 |
| Precision | bfloat16 |
| LeRobot version | 0.6.2 compatibility layer |
Use this checkpoint with the repository's lerobot-value-infer workflow to produce value, advantage, and ACP indicator fields on a compatible dataset copy. Validate field coverage, binary indicators, positive ratios, source immutability, and held-out trajectories before policy fine-tuning.
The model was trained on one SO-101 setup and task family. It may learn correlations with camera viewpoint, lighting, background, intervention behavior, or task distribution. It must be validated on held-out episodes and must not be treated as a safety controller.
This value model does not make real-robot deployment safe. Any downstream policy requires independent calibration, action-range, reset-pose, camera-mapping, and emergency-stop validation.