Downloads · 30 days
5
100% of all-time downloads
rooty2020/X-WAM-DROID
X-WAM-DROID is a robotics model from rooty2020. Use it for the robotics task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
X-WAM (Wan2.2-TI2V-5B world action model) fine-tuned on DROID with 8-D raw joint-position actions and three camera views. Same layout as the official sharinka0715/X-WAM-checkpoints release, so it drops into the upstre…
Downloads · 30 days
5
100% of all-time downloads
All-time downloads
5
Public
Repo size
22.8 GB
Likes
0
Public
Click a slice to open those files.
.pt22.8 GB · 100%
From the Hugging Face model README
X-WAM (Wan2.2-TI2V-5B world action model) fine-tuned on
DROID with 8-D raw joint-position actions and three camera views. Same layout as the
official sharinka0715/X-WAM-checkpoints
release, so it drops into the upstream X-WAM code.
| Base checkpoint | X-WAM pretrained/ (40k steps, cross-embodiment) |
| Backbone | Wan2.2-TI2V-5B DiT, UMT5-XXL text encoder, Wan2.2 VAE (stride 4×16×16) |
| Depth branch | none (use_depth: false — no extra_blocks / extra_heads) |
| Weights | bf16, full runner state dict (DiT + frozen T5 + VAE). X-WAM trains without EMA. |
| Training step | 800 (of a 20,000-step schedule) |
| Views | exterior_1_left, exterior_2_left, wrist_left at 192×320 |
| Horizon | 9 frames (frame skip 4 → 3.75 fps video), 4 actions per frame step → 32 actions at 15 Hz |
DROID raw joint positions [joint_0 … joint_6, gripper] written into X-WAM's 14-D action slots
(and 16-D proprio slots) — see action_mapping.json:
[0, 1, 2, 3, 4, 5, 7, 6] (gripper goes to slot 6)y = clip(2·(x − q01)/(q99 − q01) − 1, −1, 1), then the gripper channel is
negated, so after normalization +1 = open, −1 = closed (X-WAM convention; raw DROID is
0 = open, 1 = closed)q01 / q99 are in both config.yaml and action_mapping.jsonDecode predictions by undoing those steps in reverse order.
config.yaml training config (+ action_num, normalization stats)
action_mapping.json DROID 8-D <-> X-WAM 14-D mapping and normalization
checkpoints/last.ckpt/checkpoint/mp_rank_00_model_states.pt {"module": state_dict, "global_step", "epoch"}
config.yaml is the run's own config with one addition: action_num: 4, which the training
dataset sets at runtime and the runner needs to build the model.
hf download rooty2020/X-WAM-DROID --local-dir checkpoints/droid
Then point the upstream X-WAM scripts at it like any official checkpoint, e.g.
evaluation/policy_server.py, which loads
checkpoints/last.ckpt/checkpoint/mp_rank_00_model_states.pt with a strict
load_state_dict(ckpt["module"]). Build the runner from this config.yaml
(use_depth: false) — the official pretrained config has a depth branch and will not match.
Converted from the Lightning FSDP sharded checkpoint last.ckpt (step 800) with
xwam-droid/scripts/export_hf.py: the DiT tensors are cast fp32 → bf16 and keyed at runner level
(model.*); the frozen text_encoder.* / vae.* tensors are copied from the official pretrained
release, since they are never trained.
Apache 2.0, following X-WAM and Wan2.2. Training data comes from DROID; its terms apply to the data.