Downloads · 30 days
44
100% of all-time downloads
fan91/ImageWAM-FLUX.2-4B
ImageWAM-FLUX.2-4B is a image-to-image model from fan91. Use it when you need one image transformed into another. It is set up for imagewam. The card lists the license as other.
This is an ImageWAM FLUX.2 Klein 4B checkpoint pretrained on InternData-A1 with a 16-dimensional end-effector action representation. It is intended as an initialization model for downstream robot-policy fine-tuning.
Downloads · 30 days
44
100% of all-time downloads
All-time downloads
44
Public
Parameters
4.6B
9.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.4 GB · 100%
How the weights are stored.
BF164.5B · 98%
From the Hugging Face model README
This is an ImageWAM FLUX.2 Klein 4B checkpoint pretrained on InternData-A1 with a 16-dimensional end-effector action representation. It is intended as an initialization model for downstream robot-policy fine-tuning.
ImageWAM jointly predicts a future image and a future action chunk from camera observations, a language instruction, and robot proprioception. Its FLUX.2 video expert and ActionDiT action expert interact through a Mixture-of-Transformers (MoT) architecture.
| Item | Value |
|---|---|
| Video expert | FLUX.2 Klein 4B |
| Action expert | ActionDiT |
| Pretraining data | InternRobotics/InternData-A1 |
| Action dimension | 16 |
| Proprioception dimension | 16 |
| Maximum action horizon | 64 |
| Text encoder | Qwen/Qwen3-4B |
| Text feature dimension | 7,680 |
| Precision | bfloat16 |
model.safetensors contains:
Qwen3 weights, dataset normalization statistics, and optimizer state are not included.
@misc{zhang2026imagewam,
title={ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?},
author={Yuyang Zhang and Wenyao Zhang and Zekun Qi and He Zhang and Haitao Lin and Jingbo Zhang and Yao Mu and Xiaokang Yang and Wenjun Zeng and Xin Jin},
year={2026},
eprint={2606.19531},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.19531}
}