Downloads · 30 days
16
100% of all-time downloads
rooty2020/Cosmos3-ours-DROID
Cosmos3-ours-DROID is a robotics model from rooty2020. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for cosmos. The card lists the license as other.
A Cosmos3-Nano video-and-action policy post-trained on DROID with the Omni-4D multi-view stack: all camera views of an episode are packed along a view axis into one sequence, and the video side is trained with diffusi…
Downloads · 30 days
16
100% of all-time downloads
All-time downloads
16
Public
Parameters
15.2B
31.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors31.5 GB · 100%
From the Hugging Face model README
A Cosmos3-Nano video-and-action policy post-trained on DROID with the Omni-4D
multi-view stack: all camera views of an episode are packed along a view axis into one
sequence, and the video side is trained with diffusion forcing. The action stream uses the
backbone's inline action pathway (action2llm / llm2action / action_modality_embed);
there is no external action expert here (action_expert: null).
| Base model | nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert) |
| Architecture | cosmos3_omni, unified_3d_mrope, Omni-4D multi-view packing, two-way joint attention |
| Parameters | 15.17 B (incl. the Qwen3-VL ViT tower) |
| Weights | EMA weights, bf16 |
| Training iteration | 250 (early checkpoint of a 5000-step schedule) |
| Warm start | Cosmos3-Nano-Policy-DROID |
| Action space | 64-dim zero-padded, domain-aware I/O, 32 embodiment domains |
require_all_views, no view cap), action chunk 16, history latents {0,1,2},
history actions and state on, 92k tokens per packed sample.A standard consolidated Cosmos checkpoint (config.json + sharded model*.safetensors +
checkpoint.json), the same layout nvidia/Cosmos3-Nano ships:
hf download rooty2020/Cosmos3-ours-DROID --local-dir ./Cosmos3-ours-DROID
torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
-i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-ours-DROID
Multi-view rollout and action decoding expect the Omni-4D packer; training_config.yaml
carries the full training configuration of the source run.
Exported from a PyTorch Distributed Checkpoint (iter 250) with
python -m cosmos_framework.scripts.export_model --use-ema-weights. The ViT tower is not in
the training checkpoint and is taken from Qwen/Qwen3-VL-8B-Instruct at the revision pinned
by the Cosmos framework.
Derived from nvidia/Cosmos3-Nano and governed by the
NVIDIA Open Model License. Training data comes
from DROID; its terms apply to the data. The usual
caveats about generated video and learned policies (no guarantee of physical accuracy, not
for safety-critical control) apply.