Downloads · 30 days
199
29% of all-time downloads
nvidia/Cosmos-Policy-ALOHA-Predict2-2B
Cosmos-Policy-ALOHA-Predict2-2B is a machine learning model from nvidia. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Cosmos Policy | Code | White Paper | Website
Downloads · 30 days
199
29% of all-time downloads
All-time downloads
691
Public
Repo size
3.9 GB
Likes
8
Public
Click a slice to open those files.
.pt3.9 GB · 100%
From the Hugging Face model README
Cosmos Policy | Code | White Paper | Website
Cosmos-Policy-ALOHA-Predict2-2B is a 2B-parameter bimanual robot manipulation policy model fine-tuned from the NVIDIA Cosmos-Predict2-2B-Video2World video foundation model. This model achieves a 93.6% average completion rate across four challenging real-world bimanual manipulation tasks on the ALOHA 2 robot platform.
Key features:
Use cases:
This model is for research and development only.
Model Developer: NVIDIA
Cosmos Policy models include the following:
This model is released under the NVIDIA One-Way Noncommercial License (NSCLv1). For a custom license, please contact [email protected].
Under the NVIDIA One-Way Noncommercial License (NSCLv1), NVIDIA confirms:
Global
Physical AI: Bimanual robot manipulation and control in real-world environments, encompassing contact-rich manipulation and imitation learning from human demonstrations.
GitHub [01/22/2026] via https://github.com/nvlabs/cosmos-policy
Hugging Face [01/22/2026] via https://huggingface.co/collections/nvidia/cosmos-policy
Architecture Type: A diffusion transformer with latent video diffusion, fine-tuned from Cosmos-Predict2-2B-Video2World.
Network Architecture: The model uses the same architecture as the base Cosmos-Predict2-2B model (a diffusion transformer with latent video diffusion).
Key adaptation: Actions, proprioceptive states, and values are encoded as latent frames and injected directly into the video model's latent diffusion sequence, enabling the model to generate these modalities alongside predicted future images.
Number of model parameters:
2B (inherited from base model)
Input Type(s): Text + Multi-view Images + Proprioceptive State
Input Format(s):
Input Parameters:
Other Properties Related to Input:
Output Type(s): Action Sequence + Future State Predictions + Value Estimate
Output Format:
Output Parameters:
Other Properties Related to Output:
Note on future predictions: The future state images and value predictions generated by this base policy checkpoint are primarily for visualization and interpretability purposes. For model-based planning with these predictions, please additionally use the separate Cosmos-Policy-ALOHA-Planning-Model-Predict2-2B checkpoint as the world model and value function.
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Runtime Engine(s):
Supported Hardware Microarchitecture Compatibility:
Note: We have only tested doing inference with BF16 precision.
Operating System(s):
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
Hardware Compatibility Warning: This model was trained on a specific ALOHA 2 robot setup with particular hardware characteristics. Differences between our robot setup and downstream users' hardware setups (including calibration, joint limits, camera positioning, gripper mechanics, etc.) may significantly impact performance. Users must exercise caution during deployment.
Control Frequency: This policy must be used with a 25 Hz controller for satisfactory performance (not the original 50 Hz ALOHA control frequency). The reduced frequency was used during data collection and training.
Real-World Deployment: This model operates real robotic hardware. Always ensure that proper safety measures are in place. On the first deployment of this checkpoint, we highly recommend measuring the difference in the current robot state and the next commanded robot state (e.g., difference between current joint angles and predicted actions, which represent target joint angles) and aborting policy execution if the difference is large.
See Cosmos Policy GitHub for details.
Data Collection Method:
Labeling Method:
Training Data: ALOHA-Cosmos-Policy dataset
Training Configuration:
model-480p-16fps.pt)Training Objective: The model is trained with a hybrid log-normal-uniform noise distribution (modified from the base model's log-normal distribution; see paper for details) to improve action prediction accuracy. Training batches are split 50/25/25 for policy, world model, and value function objectives, respectively.
Data Collection Method: Not Applicable
Labeling Method: Not Applicable
Properties: Not Applicable - We use the real-world ALOHA 2 robot platform for direct evaluations.
Test Hardware: H100, A100
See Cosmos Policy GitHub for details.
Inference with base Cosmos Policy only (i.e., no model-based planning):
| Task | Score |
|---|---|
| put X on plate | 100.0 |
| fold shirt | 99.5 |
| put candies in bowl | 89.6 |
| put candy in ziploc bag | 85.4 |
| Average | 93.6 |
Scores represent average percent completion across 101 trials total (including both in-distribution and out-of-distribution test conditions).
Comparison with baselines:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.
Please report security vulnerabilities or NVIDIA AI Concerns here.
If you use this model, please cite the Cosmos Policy paper:
(Cosmos Policy BibTeX citation coming soon!)