Downloads · 30 days
138
3% of all-time downloads
robotics-diffusion-transformer
robotics-diffusion-transformer/RDT2-FM
RDT2-FM is a robotics model from robotics-diffusion-transformer. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Using a flow-matching objective, RDT2-FM delivering lower inference latency while preserving strong instruction following and cross-embodiment generalization on UMI-style bimanual setups. Concretely, This repository c…
Downloads · 30 days
138
3% of all-time downloads
All-time downloads
5K
Public
Repo size
2 GB
Likes
10
Public
Click a slice to open those files.
.bin975 MB · 100%
From the Hugging Face model README
RDT2-FM builds on a vision-language backbone (RDT2-VQ) and predicts short-horizon relative action chunks through an action expert that integrates an improved RDT architecture with a flow-matching objective. By leveraging flow matching, RDT2-FM achieves lower inference latency while maintaining strong instruction following and cross-embodiment generalization on UMI-style bimanual setups. This repository specifically provides the action expert component of RDT2-FM.
Home - Github - Discord - Paper
20-D per step = right (10) + left (10):
Output tensor shape: (T=24, D=20), relative deltas, float32.
Approximate single-GPU requirements:
| Mode | RAM | VRAM | Example GPU |
|---|---|---|---|
| Inference (FM head + VLM) | ≥ 32 GB | ~ 16 GB | RTX 4090 |
| Fine-tuning FM head | – | ~ 16 GB | RTX 4090 |
For deployment on real robots, follow your platform’s end-effector + camera choices and perform hardware setup & calibration (camera stand/pose, flange, etc.) before running closed-loop policies.
Tested OS: Ubuntu 24.04.
# Run under root directory of RDT2 GitHub Repo: https://github.com/thu-ml/RDT2/tree/main?tab=readme-ov-file#1-important-hard-ware-set-up-and-calibration
import yaml
from models.rdt_inferencer import RDTInferencer
with open("configs/rdt/post_train.yaml", "r") as f:
model_config = yaml.safe_load(f)
model = RDTInferencer(
config=model_config,
pretrained_path="robotics-diffusion-transformer/RDT2-FM",
# TODO: modify `normalizer_path` to your own downloaded normalizer path
# download from http://ml.cs.tsinghua.edu.cn/~lingxuan/rdt2/umi_normalizer_wo_downsample_indentity_rot.pt
normalizer_path="umi_normalizer_wo_downsample_indentity_rot.pt",
pretrained_vision_language_model_name_or_path="robotics-diffusion-transformer/RDT2-VQ", # use RDT2-VQ as the VLM backbone
device="cuda:0",
dtype=torch.bfloat16,
)
result = model.step(
observations={
'images': {
# 'exterior_rs': np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8),
'left_stereo': ..., # left arm RGB image in np.ndarray of shape (384, 384, 3) with dtype=np.uint8
'right_stereo': ..., # right arm RGB image in np.ndarray of shape (384, 384, 3) with dtype=np.uint8
},
# use zero input current state for currently
# preserve input interface for future fine-tuning
'state': np.zeros(model_config["common"]["state_dim"]).astype(np.float32)
},
instruction=instruction # Language instruction
# We suggest using Instruction in format "verb + object" with Capitalized First Letter and trailing period
)
# relative action chunk in np.ndarray of shape (24, 20) with dtype=np.float32
# with the same format as RDT2-VQ
action_chunk = result.detach().cpu().numpy()
# rescale gripper width from [0, 0.088] to [0, 0.1]
for robot_idx in range(2):
action_chunk[:, robot_idx * 10 + 9] = action_chunk[:, robot_idx * 10 + 9] / 0.088 * 0.1
For guides on installation and fine-tuning, please refer to the official GitHub repository.
bfloat16 for training and inference.bfloat16 by default (Qwen2.5-VL practices).Intended uses
Limitations
Safety & responsible use
| Symptom | Likely cause | Suggested fix |
|---|---|---|
| Drifting / unstable gripper widths | Scale mismatch | Apply LinearNormalizer; rescale widths ([0,0.088] → [0,0.1]). |
| Poor instruction following | Prompt format / backbone config | Use “Verb + Object.”; ensure backbone is loaded on same device. |
@article{liu2026rdt2,
title={RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization},
author={Liu, Songming and Li, Bangguo and Ma, Kai and Wu, Lingxuan and Tan, Hengkai and Ouyang, Xiao and Su, Hang and Zhu, Jun},
journal={arXiv preprint arXiv:2602.03310},
year={2026}
}