Downloads · 30 days
1.4K
65% of all-time downloads
nvidia/Cosmos-H-Dreams
Cosmos-H-Dreams is a image-to-video model from nvidia. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for nv-medtech. The card lists the license as other.
<div align="center" <div align="center" <a href="https://github.com/isaac-for-healthcare/Cosmos-H-Dreams" <img src="https://img.shields.io/badge/GitHub-Cosmos--H--Dreams-grey?logo=GitHub" alt="Runtime Badge" </a <a hr…
Downloads · 30 days
1.4K
65% of all-time downloads
All-time downloads
2.1K
Public
Repo size
4.4 GB
Likes
16
Public
Click a slice to open those files.
.pt4.4 GB · 100%
From the Hugging Face model README
Cosmos-H-Dreams is a real-time, action-conditioned generative surgical world model that lets a human operator or a learned surgical-robotics policy act inside a synthesized surgical scene and observe the interactions live. Given a surgical context frame and a stream of robot kinematic actions, the model autoregressively generates the resulting future video in short blocks, streaming interactively on a single GPU.
Unlike the bidirectional, offline Cosmos-H-Surgical-Simulator, Cosmos-H-Dreams is a causal, few-step self-forcing distilled student: it is distilled from a bidirectional teacher into a streaming model that responds immediately to actions, turning a passive video generator into a controllable surgical simulator. The released checkpoint specializes the model to da Vinci Research Kit (dVRK) tabletop suturing.
The model is intended for real-time surgical-skills rehearsal, interactive demonstration, closed-loop evaluation of surgical robotics policies, and synthetic data generation.
The released model is derived from the public NVIDIA Cosmos-Predict2.5-2B world foundation model for physical AI, and its teacher is warm-started from Cosmos-H-Surgical-Simulator (the Open-H 44D action-conditioned checkpoint).
This model is ready for commercial or non-commercial use.
Use of this model is governed by the NVIDIA Open Model License Agreement.
Global
Primarily intended for surgical robotics researchers, healthcare AI developers, academic institutions, and surgical robotics companies exploring interactive surgical simulation, surgeon/trainee rehearsal, closed-loop surgical policy evaluation, and synthetic data generation. Because it is interactive, the same served model can be driven live by a human (browser keyboard or Meta Quest headset) or by a learned policy in a closed-loop evaluation harness.
Ali, A., Bai, J., Bala, M., Balaji, Y., Blakeman, A., Cai, T., Cao, J., Cao, T., Cha, E., Chao, Y.-W., Chattopadhyay, P., Chen, M., Chen, Y., Cheng, S., Cui, Y., Diamond, J., Ding, Y., Fan, J., Fan, L., Feng, L., Ferroni, F., Fidler, S., Fu, X., Gao, R., Ge, Y., Gu, J., … Zhu, Y. (2025). World Simulation with Video Foundation Models for Physical AI (arXiv:2511.00062) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.00062
Link to Cosmos’ nvidia/Cosmos-Predict2.5-2B-Video2World Model Card
Link to nvidia/Cosmos-H-Surgical-Simulator Model Card (the teacher warm-start checkpoint).
Architecture Type: Diffusion Transformer Network Architecture: Latent video diffusion transformer (DiT-style denoiser) with cross-attention conditioning, distilled into a causal, few-step autoregressive student with a streaming key/value (KV) cache.
This model was developed based on Cosmos-Predict2.5-2B-Video2World.
Cosmos-H-Dreams extends Cosmos-Predict2.5-2B-Video2World, a 2B-parameter diffusion transformer for video generation in latent space (Wan2.1 video tokenizer, spatial compression 8×, temporal compression 4×, 16 latent channels; hidden width 2048, 28 blocks, 16 heads, 2×2 spatial patchification, rotary position embeddings, AdaLN-LoRA modulation, Cosmos-Reason text encoder cross-attention). It incorporates two small MLPs that condition the model on kinematic actions through the timestep/AdaLN modulation pathway.
The model consumes a unified 44-dimensional action vector; each latent frame folds the 4 action steps it is responsible for into a 4 × 44 = 176-dimensional input. The unified action space lets a subset of the 44 dimensions carry an embodiment's native content while the remainder are zero-padded, which keeps the weights embodiment- and horizon-invariant. The released dVRK tabletop checkpoint uses a 20-dimensional dual-arm content vector (per patient-side manipulator: a 3D translation delta, a 6D continuous rotation, and a gripper value) zero-padded to 44D.
The model is built in two stages from the same backbone:
.npy file is required.Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Runtime Engine(s):
torch.compile + CUDA Graphs, few-step diffusion, lightweight LightTAE decoder, and the keyboard/Quest serving stack).Supported Hardware Microarchitecture Compatibility:
Note: Only BF16 (Brain Floating Point 16) precision is tested. Other precisions like FP16 or FP32 are not officially supported.
Preferred/Supported Operating System(s): Linux (We have not tested on other operating systems.)
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
v1.0 Real-time self-forcing distilled student (dVRK tabletop suturing). A causal, few-step (2–4 step) streaming student distilled from a bidirectional teacher that is warm-started from the Cosmos-H-Surgical-Simulator Open-H 44D checkpoint and fine-tuned on the dVRK tabletop suturing mixture (a subset of the Open-H Embodiment dataset).
Developers may integrate the model into an AI evaluation or rehearsal system by providing a context video frame along with a stream of kinematic actions, either from a live control surface (keyboard or Meta Quest) or from a surgical policy model such as GR00T-H and receiving streamed video of the resulting scene.
The teacher is warm-started from the Open-H 44D pre-training checkpoint (Cosmos-H-Surgical-Simulator) and then fine-tuned/distilled on the target regime. The released student targets dVRK tabletop suturing.
Dataset: Open-H-Embodiment community-generated dataset (pre-training) and the dVRK tabletop suturing dataset (post-training).
Key performance: The model can generate physically plausible surgical robotic videos in real time given an initial frame and a live kinematic action stream, either from a surgeon/trainee (keyboard or Meta Quest) or from a surgical robotic policy model such as GR00T-H.
Acceleration Engine: PyTorch, Transformer Engine, torch.compile + CUDA Graphs (via FlashDreams).
Test Hardware: Real-time interactive serving validated on a single NVIDIA RTX PRO 6000 (edge target); also runs on NVIDIA Ampere (A100) and NVIDIA Hopper (H100)-class GPUs.
Usage:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please make sure you have proper rights and permissions for all input image and video content; if image or video includes people, personal health information, or intellectual property, the image or video generated will not blur or maintain proportions of image subjects included.
For more detailed information on ethical considerations for this model, please see the Model Card++ for Bias, Explainability, Safety & Security, and Privacy Subcards.
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.