Downloads · 30 days
0
armanakbari4/imagewam-ur3-3task
imagewam-ur3-3task is a robotics model from armanakbari4. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Real-world fine-tune of ImageWAM (FLUX.2 [klein] base 4B editing DiT + ActionDiT action expert) on a bimanual dual-arm UR3, trained jointly on three tasks:
Downloads · 30 days
0
Access
Public
Updated Aug 21, 2026
Repo size
45.2 GB
Likes
0
Public
Click a slice to open those files.
.pt45.2 GB · 100%
From the Hugging Face model README
Real-world fine-tune of ImageWAM (FLUX.2 [klein] base 4B editing DiT + ActionDiT action expert) on a bimanual dual-arm UR3, trained jointly on three tasks:
| task | episodes | frames | instruction |
|---|---|---|---|
blue_basket | 100 | 29,653 | "put the medicine then the measuring tape inside the blue basket" |
drawer | 100 | 32,982 | "open the drawer, put the white box inside the drawer then close the drawer" |
stacking_cubes | 100 | 37,042 | "put the green cube on top of the black cube and put the red cube on top of the green cube" |
300 episodes / 99,677 frames total, 15 fps. Initialized from
yuyangalin/ImageWAM-FLUX.2-4B-InternData-A1-EE
(InternData-A1 pretrain, step 60k) — loaded with 0 missing / 0 unexpected keys.
| File | held-out action_l1 |
|---|---|
ur3_3task_ee16_step2000.pt | 0.0299 |
ur3_3task_ee16_step3000.pt | 0.0304 |
ur3_3task_ee16_step4000.pt | 0.0281 (nominal best) |
ur3_3task_ee16_step7000.pt | 0.0295 |
ur3_3task_ee16_step10000.pt | 0.0341 (final) |
ur3_3task_ee16_dataset_stats.json | z-score stats — required for inference |
train_config.yaml | resolved training config |
Each .pt holds only the trained parts (mot ≈ 8.2 B params + proprio_encoder), ~9.0 GB.
FLUX.2 klein-base-4B base weights and autoencoder must be prepared separately.
lr 2.5e-5 cosine w/ 5% warmup, AdamW(0.9, 0.95), wd 1e-2, grad-clip 1.0, bf16,
DeepSpeed ZeRO-1, global batch 192 (8/GPU × 6 GPUs × 4 grad-accum), 10,000 steps
(~19 epochs), ~9.6 h on 6×H100. num_frames=17, action_video_freq_ratio=1 →
16-step action horizon, endpoint_frames_only=true. 14-dim UR3 joint action/state is
padded to the pretrain's 16D space at dims 7 and 15 (ee16); nothing is converted to
end-effector poses.
compact_288x256): camera_top 192×256 on top, the two wrists
96×128 side-by-side below, order fixed [top, left, right], pixels normalized to (−1, 1).concat(x[..., 0:7], x[..., 8:15]) — before
sending to the controller.ur3_3task_ee16_dataset_stats.json for denormalization; these stats are computed
on this dataset, not the pretrain's.Train loss_action fell 21× over the run (0.0917 → 0.0043), but held-out action_l1
barely moved: by thirds of training 0.0344 → 0.0325 → 0.0306, with a regression slope of
only −0.00053 per 1k steps (t = −2.14) across 20 evals. Per-eval noise (sd 0.0034 over
32 clips) is larger than the gap between any two of the checkpoints above, so these five are
statistically indistinguishable and the ranking should not be trusted. Step 4,000's 0.0281 sits
between 0.0323 and 0.0362 at the adjacent evals.
Checkpoints exist only at multiples of 1,000 (save_every: 1000); evals ran every 500, so
some of the best-scoring evals (e.g. 0.0248 at step 7,500) have no corresponding checkpoint.
Offline action_l1 has not been validated against real-robot success rate on this setup.
Select by real-robot success, not by this metric.
Note also that the last 3,000 steps bought nothing measurable — 7k would have sufficed.
Built on ImageWAM:
@misc{zhang2026imagewam,
title={ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?},
author={Yuyang Zhang and Wenyao Zhang and Zekun Qi and He Zhang and Haitao Lin and Jingbo Zhang and Yao Mu and Xiaokang Yang and Wenjun Zeng and Xin Jin},
year={2026},
eprint={2606.19531},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2606.19531},
}