Downloads · 30 days
0
Kasra99/groot_dex_warehouse_ae_2
groot_dex_warehouse_ae_2 is a robotics model from Kasra99. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for lerobot. The card lists the license as other.
GR00T N1.7 action-expert fine-tune on Kasra99/dex-warehouse, a teleoperated warehouse pick-and-place dataset recorded on a Dexmate Vega 1 Pro mobile manipulator.
Downloads · 30 days
0
Access
Public
Updated Aug 29, 2026
Repo size
62.9 GB
Likes
0
Public
Click a slice to open those files.
.safetensors62.9 GB · 100%
From the Hugging Face model README
GR00T N1.7 action-expert fine-tune on Kasra99/dex-warehouse,
a teleoperated warehouse pick-and-place dataset recorded on a Dexmate Vega 1 Pro mobile manipulator.
ae = action expert: the Cosmos-Reason2 / Qwen3-VL backbone (LLM + vision tower) is frozen; the
flow-matching action head is trained.
| Base weights | nvidia/GR00T-N1.7-3B |
| Embodiment tag | new_embodiment (projector slot 10) |
| Trainable | projector 327M + DiT 1.09B + vlln/vl-attn 201M = 1.62B / 3.14B (51.5%) |
| Frozen | LLM 1.12B + vision tower 407M |
| Steps / batch | 45,000 / 32 (5.9 epochs over 245,541 frames) |
| Optimizer | AdamW, lr 1e-4, cosine, warmup |
| Chunk / action steps | 40 / 40 (N1.7 native horizon) |
| EMA | constant decay 0.99 (weights published are the EMA weights) |
| Augmentation | photometric jitter — brightness, contrast, saturation, hue, sharpness; ≤3 per frame |
| Validation split | none — all 213 episodes used for training |
new_embodiment maps to embodiment id 10, which is unused in NVIDIA's pretraining (absent from both
embodiment_id.json and statistics.json). Its category-specific projector is therefore randomly
initialised and trained from scratch on this robot; normalization statistics come from the dataset.
Three cameras, renamed to GR00T/π-style keys:
observation.images.base_0_rgb (head camera)
observation.images.left_wrist_0_rgb
observation.images.right_wrist_0_rgb
20-dimensional state and action, in this order:
0 arm_center_z torso lift
1-7 L_arm_j1 .. L_arm_j7 left arm joints
8-14 R_arm_j1 .. R_arm_j7 right arm joints
15 right_hand.open_close_ratio
16 right_hand.thumb_opposition_ratio
17-19 base_vx, base_vy, base_wz mobile base velocity command
The two left-hand DoF present in the raw dataset were dropped: the left hand is commanded in only
1,414 of 219,260 frames, and its stored quantile range spans ~0.01, which maps the rare 1.0 to a
normalized +199. Removing those two dimensions drops the maximum normalized action magnitude from
199 to 10.
EMA weights at five points in training:
step_009000/ step_018000/ step_027000/ step_036000/ step_045000/
There is no model at the repository root — download a step folder and load it by local path
(from_pretrained has no subfolder argument):
from huggingface_hub import snapshot_download
from lerobot.policies.groot.modeling_groot import GrootPolicy
STEP = "step_027000"
root = snapshot_download("Kasra99/groot_dex_warehouse_ae_2", allow_patterns=f"{STEP}/*")
policy = GrootPolicy.from_pretrained(f"{root}/{STEP}")
Requires lerobot >= 0.6.2 with the groot extra (pip install 'lerobot[groot]').
22 language instructions, all warehouse pick-and-place, factorising into 5 objects (banana, batman
toy, bear toy, blue bird, box) × 5 destinations (box, gaylord, table, conveyor belt, pick-only).
Coverage is uneven — box and gaylord destinations dominate, while
conveyor belt has 1,796 frames across 3 episodes. Instructions in the thin tail should not be
expected to work.
groot_dex_warehouse_aeRun 2 of the same recipe. Identical base weights, hyperparameters and seed; only the dataset
differs (+36 episodes: +17 bear toy, +18 banana, +1 batman toy — the objects the first model
handled least reliably). 13 corrupted task strings present in the raw dataset (typos such as
'pcik up the batman toy...', 'ick up the blue bird', and phrasing variants such as
'drop the box on conveyor belt') were merged into their canonical form before training, so the
language conditioning is not fragmented across near-duplicate instructions.
Step counts were chosen so the five checkpoints land on the same epochs as run 1's
(1.17 / 2.35 / 3.5 / 4.7 / 5.85), making the two runs directly comparable. step_027000 here is
the analogue of run 1's step_024000.