Downloads · 30 days
0
Ronaldo-GOAT/actaug-gr00t-train-eval
actaug-gr00t-train-eval is a machine learning model from Ronaldo-GOAT. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Everything needed to train GR00T-N1.5 on the action-augmentation RoboCasa datasets (actaug-prio-1024 full + its 512 subset) from the 60k base checkpoint, and to evaluate the resulting checkpoints on the 160-episode ex…
Downloads · 30 days
0
Access
Public
Updated Sep 30, 2026
Repo size
7.6 GB
Likes
0
Public
Click a slice to open those files.
.safetensors7.6 GB · 100%
From the Hugging Face model README
Everything needed to train GR00T-N1.5 on the action-augmentation RoboCasa datasets
(actaug-prio-1024 full + its 512 subset) from the 60k base checkpoint, and to
evaluate the resulting checkpoints on the 160-episode exact-replay benchmark.
This repo is self-contained: it ships the training code (with our save/stop knobs), the eval stack (with the porting fixes), the base checkpoint, the subset builder, and this protocol doc. Large external artifacts (the training datasets, the eval replay bank) are on HF and fetched by the setup steps below.
Goal. Measure how much the action-augmentation data helps GR00T on RoboCasa PickPlaceCounterToCabinet, and how that depends on dataset size and training length.
actaug-prio-1024).Why a queue? Training emits checkpoints over time (4k appears first, then 4.5k, 5k, …). Each eval takes ~2–3 h on one GPU. If we waited for training to finish and then evaluated 13–14 checkpoints × 2 conditions serially, it would take days. Instead we run a small pool of eval workers that start evaluating each checkpoint the moment it lands, in parallel with training and with each other — so evals finish almost as fast as the checkpoints arrive.
A pool of 2–3 eval workers (each = 1 GPU running the 4-server/4-client 160-ep eval of one checkpoint at a time) shares one work queue across both training runs:
checkpoint-N/ that are complete (have
model.safetensors.index.json + trainer_state.json — i.e. fully flushed, not mid-write).mkdir a
<ckpt>.claim) so two workers never eval the same one.eval_one_checkpoint_local.sh on it → writes summary.json (all-160 success
rate + per-episode stage flags) into that checkpoint's eval dir."Checkpoint-wise FCFS, not run-wise" means the two training runs' checkpoints are interleaved into one queue and taken as they become ready — a worker takes whichever eligible checkpoint is next by the ordering above, regardless of which run produced it.
# 1. env: mygr00t (train + serve), robocasa (eval client) — see §1
# 2. base checkpoint: base60k/ (shipped here) == mlnha/gr00t-n15-robocasa-base60k
# 3. datasets: Ronaldo-GOAT/actaug-prio-1024 (full 1024); build the 512 subset with scripts/build_512_subset.py
# 4. TRAIN (2 GPU, premium, 30k-schedule but STOP at 10k, save 4k->10k every 500):
sbatch -p sjw_alinlab_premium --wckey=project-short-name:others --export=ALL scripts/train_actaug.sbatch
# 5. EVAL each checkpoint (1 GPU, 4 workers, exact160 bank, POLICY_SEED=12345):
MODEL_PATH=<ckpt> ROOT=<out> bash code/eval/eval_one_checkpoint_local.sh
Two conda envs (GR00T serving and the RoboCasa sim client have incompatible deps):
mygr00t — trains and serves the GR00T policy. PYTHONPATH=code/myGR00T. Has torch+CUDA+gr00t.robocasa — runs the eval client (drives the RoboCasa sim). Has robocasa==1.0.0 + a
vendored robosuite with load_model_on_init. The eval client also imports gr00t.eval.wrappers,
so code/myGR00T must be on its PYTHONPATH (the runner sets this).RoboCasa assets must be writable (robosuite writes a temp processed XML into each object's asset
dir). If the assets tree is a read-only mount, make a writable symlink farm
(cp -as <ro-assets>/. <writable>/) and point robocasa's models/assets symlink at it.
base60k/ == mlnha/gr00t-n15-robocasa-base60k (GR00T-N1.5, RoboCasa PnPCounterToCab, global
step 60000). All fine-tunes start from here (--base-model-path base60k).
Ronaldo-GOAT/actaug-prio-1024 (HF dataset) — 1024 eps = 128 per object × 8 GR00T objects
(donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, teapot_7),
3 views (256×256), 20 fps. meta/modality.json is already correct for single_panda_gripper
(video keys left_view/right_view/wrist_view → observation.images.robot0_agentview_{left,right} /
robot0_eye_in_hand; annotation human.action.task_description; rotation_type: quaternion on
base_rotation + eef_rotation_relative). No modality fix needed.index/actaug_prio_512/episode_index.jsonl
(first 64/object, "sample A"; actaug_prio_512b is the disjoint last-64/object "sample B"):
python scripts/build_512_subset.py --full <actaug_prio_1024> --out <actaug_prio_512> --index-name actaug_prio_512
This makes a lightweight VIEW (symlinks data/+videos/, subset meta/episodes.jsonl) — no parquet rewrite.Trainer: code/myGR00T/scripts/gr00t_finetune.py (env mygr00t, PYTHONPATH=code/myGR00T, cwd there).
Config = the "30k config, terminate at 10k":
| flag | value | why |
|---|---|---|
--data-config | single_panda_gripper | matches base-60k embodiment |
--embodiment-tag | new_embodiment | |
--backbone-model-type / --backbone-select-layer | eagle / 12 | |
--base-model-path | base60k | warm start |
--num-gpus / --batch-size | 2 / 32 | global batch 64 (32×2, DDP). VRAM is set by per-device batch (bs32 fits comfortably; literal bs64 on 1 GPU ≈ 70 GB, bs16 ≈ 26 GB). |
--max-steps | 30000 | schedule horizon — warmup=0.05×30000≈1500, cosine over 30k, so at step N the LR matches a full 30k run (NOT a run that decays to ~0 by N). |
--stop-at | 10000 | terminate here while keeping the 30k schedule (custom SaveAtStepsCallback). |
--save-at | 4000,4500,…,10000 (full); 2000,4000,4500,…,10000 (512) | explicit save steps; periodic saving disabled. |
Two knobs we added to the stock trainer (--save-at, --stop-at) via SaveAtStepsCallback
— saves exactly at listed steps and stops at stop_at without shortening the LR schedule.
Submit (per §0). The sbatch (scripts/train_actaug.sbatch) waits for the dataset to be fully present,
then trains; on preemption it auto-detects the newest checkpoint and passes --resume (the
save-at checkpoints are full — optimizer+scheduler+trainer_state — so the cosine continues). Use
sjw_alinlab_premium (non-preemptible) to avoid requeue churn.
Cluster submit-filter rules (this cluster): pass --wckey=project-short-name:others; do NOT pass
--cpus-per-task; export MODEL_OUTPUT_DIR starting with /rlwrld-unified-checkpoints/jonghoon/;
comma-valued env vars (like --save-at) must go through --export=ALL on an exported shell var, NOT
inline in --export=ALL,SAVE_AT=... (commas are the --export delimiter → truncation).
Exact-state replay on PickPlaceCounterToCabinet, 160 episodes = 8 objects × 20, restoring
the exact recorded MuJoCo scene per episode. Bank:
Ronaldo-GOAT/transfer :: actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506
(its 8 objects are exactly this dataset's training objects).
Per checkpoint (1 GPU, 4 servers + 4 clients, ~2–3 h):
MODEL_PATH=<step_dir> ROOT=<eval_out> CUDA_DEV=0 \
bash code/eval/eval_one_checkpoint_local.sh
Fixed protocol: POLICY_SEED=12345 (server RNG + per-step policy_seed+ep_idx*1000003+step),
SEED_BASE=42, ACTION_HORIZON=16, N_EPISODES=160. 4 workers = length of GPUS list
(GPUS=0,0,0,0 packs 4 on one GPU). Writes summary.json + per-episode stage flags
(grasped/grasped_strict/lifted/in_cab). Baseline: base-60k = 11/160 (6.9%) on this set.
FCFS eval queue (evaluate checkpoints as training produces them, concurrently): run 2–3 of the
above as independent single-GPU jobs, each claiming the next un-evaluated checkpoint, integer-k
first (4k,5k,6k,… before 4.5k,5.5k,…). POLICY_SEED is seed-sensitive — don't over-read small
single-seed differences.
The verified eval was captured on the original NVIDIA cluster; the replay bank bakes absolute
/lp-dev/... asset paths. To run elsewhere we:
model_xml_gz/state_npz/ep_meta_pickle to the local bank,edit_model_xml() on the loaded scene XML (repaths recorded assets to the local
install — mirrors robocasa's own playback_dataset.reset_to),--generative_textures (GENERATIVE_TEXTURES=0) — it only affects the thrown-away first
reset; the real scene comes from the recorded XML, so results are unaffected.
These are porting-only; the replayed scene and policy I/O are unchanged. GPU architecture is a
controlled variable — compare checkpoints only on the same GPU model.sjw_alinlab_premium.README.md this protocol
base60k/ the 60k base checkpoint (== mlnha/gr00t-n15-robocasa-base60k)
code/myGR00T/ training + serving code (adds --save-at / --stop-at)
code/eval/ eval stack (exact-replay runner + client, with porting fixes)
eval_groot15_exact_replay.sh multi-worker eval launcher (server + clients)
eval_robocasa_replay_state_grasp.py robocasa replay client (+ stage flags)
eval_one_checkpoint_local.sh local wrapper: 1 GPU / 4 workers, all env overrides set
EVALUATION.md upstream eval notes
scripts/train_actaug.sbatch 2-GPU training job (dataset guard + resume-on-preempt)
scripts/build_512_subset.py build the 512 view from the full dataset's index list