Downloads · 30 days
6.3K
8% of all-time downloads
nvidia/Cosmos-H-Surgical
Cosmos-H-Surgical is a image-to-video model from nvidia. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for cosmos-h-surgical. The card lists the license as other.
<div align="left" <a href="https://github.com/NVIDIA-Medtech/Cosmos-H-Surgical" <img src="https://img.shields.io/badge/GitHub-grey?logo=GitHub" alt="GitHub Badge" </a <a href="https://arxiv.org/abs/2512.23162" <img sr…
Downloads · 30 days
6.3K
8% of all-time downloads
All-time downloads
76.6K
Public
Parameters
15.2B
108 GB on disk
Likes
29
Trending 3
Click a slice to open those files.
.safetensors61.9 GB · 100%
From the Hugging Face model README
Cosmos-H-Surgical v0.3.1 provides two checkpoints built on NVIDIA Cosmos3-Nano:
| Model key | Repository path | Use |
|---|---|---|
Cosmos-H-Surgical | Repository root | Default model for video prediction and 50-step surgical transfer |
Cosmos-H-Surgical-Transfer-DMD2-4Step | dmd2-transfer-480p-4step/ | Four-denoiser-call DMD2 student for surgical transfer |
The base Cosmos-H-Surgical model remains the default when
--checkpoint-path is omitted. The DMD2 checkpoint is an explicitly selected,
student-only safetensors export; it does not contain optimizer, scheduler,
teacher, fake-score, trainer, or other DCP training state.
The supported workflows are:
Each transfer input specification activates exactly one control modality.
This model is ready for commercial and non-commercial use.
The model weights and accompanying Cosmos-H-Surgical source code are provided under OpenMDW-1.1. The official license source is available from the OpenMDW repository. Third-party software and upstream model components remain subject to their own license terms.
Expected users include medical researchers, surgical robotics developers, AI developers, and healthcare institutions working on synthetic surgical video, simulation-to-real transfer, physical AI, and downstream robotics research.
These models are not intended for clinical diagnosis or autonomous clinical decision-making. Validate model behavior with use-case-specific data before deployment. Users must have appropriate rights and permissions for all input images and videos.
| Property | Base checkpoint | DMD2 transfer checkpoint |
|---|---|---|
| Base model | NVIDIA Cosmos3-Nano | Distilled from Cosmos-H-Surgical transfer |
| Architecture | Diffusion transformer with unified language and video pathways | Student diffusion transformer with rank-16 transfer LoRA |
| Precision | BF16 | BF16 |
| Resolution | 480P, validated at 832 x 480 | 480P, validated at 832 x 480 |
| Frame rate | 16 FPS | 16 FPS |
| Output length | 93 frames | 93 frames |
| Transfer sampling | 50 denoising steps | Fixed SDE times [1.0, 0.9375, 0.8333333333333334, 0.625] |
| Guidance for validated transfer | Text 1.0, control 1.0 | Text 1.0, control 1.0; no extra CFG branch |
The DMD2 export contains seven safetensors shards, 1,092 indexed tensor keys, 15,171,557,568 parameters, and 288 LoRA tensor keys.
Clone the v0.3.1 source release and install one CUDA dependency group:
# CUDA 13, recommended
uv sync --group cu130
# CUDA 12.8
uv sync --group cu128
source .venv/bin/activate
The default base checkpoint can be run without an explicit checkpoint name:
torchrun --nproc-per-node=8 \
-m cosmos_h_surgical infer \
--parallelism-preset=latency \
--dp-shard-size=1 \
-i inputs/predict/surgical_predict.jsonl \
-o outputs/cosmos-h-surgical/predict \
--no-guardrails \
--seed=0
Select the DMD2 student explicitly for four-step transfer:
torchrun --nproc-per-node=1 \
-m cosmos_h_surgical infer \
--checkpoint-path Cosmos-H-Surgical-Transfer-DMD2-4Step \
--parallelism-preset=latency \
--dp-shard-size=1 \
--no-use-torch-compile \
-i /path/to/dmd2_transfer_input.json \
-o outputs/cosmos-h-surgical/dmd2-transfer \
--no-guardrails \
--seed=0
The DMD2 input uses model_mode: "video2video", exactly one active control,
and both guidance: 1.0 and control_guidance: 1.0. The full input schema and
examples are in the
inference documentation.
The v0.3.1 source release includes a public DMD2 training launcher, export wrapper, and tutorial. The public configuration is a recommended starting recipe rather than an exact reconstruction of the internal training run that produced this checkpoint.
The default control sampling configuration is:
control_modalities={
"edge": 1.0,
"blur": 1.0,
"depth": 1.0,
"seg": 1.0,
}
See the DMD2 distillation tutorial for data preparation, single-node and multi-node launch, resume, export, and inference instructions.
The released DMD2 student was evaluated on 1,010 videos per control modality at 832 x 480, 93 frames, and 16 FPS:
| Control | FVD |
|---|---|
| Edge | 40.3171 |
| Blur | 34.7035 |
| Depth | 37.8086 |
| Segmentation | 41.1120 |
Latency was measured on one H100 80 GB GPU using one edge-transfer sample at 832 x 480, 93 frames, and 16 FPS, with seed 0, BF16, compilation disabled, one warmup, and five measured repetitions. Values are mean +/- sample standard deviation in seconds. CUDA stage times are synchronized; generation and complete-batch times are wall-clock measurements.
| Stage | DMD2 four-step | Base 50-step | Base / DMD2 |
|---|---|---|---|
| VAE encode CUDA, two calls total | 1.107 +/- 0.000 | 1.108 +/- 0.000 | 1.00x |
| Denoising CUDA | 4.087 +/- 0.007 | 107.163 +/- 0.389 | 26.22x |
| VAE decode CUDA, one call | 1.993 +/- 0.002 | 1.994 +/- 0.002 | 1.00x |
| Generation wall time | 8.780 +/- 0.048 | 112.704 +/- 0.415 | 12.84x |
| Complete batch wall time | 10.882 +/- 0.071 | 114.940 +/- 0.471 | 10.56x |
The DMD2 run executed four denoiser calls. The base run executed 100 calls (50 steps with two classifier-free guidance evaluations per step). The DMD2 run did not execute an additional CFG branch.
The release validation also checks all four public transfer controls, verifies the exported checkpoint inventory and checksums, and compares a fixed-seed exported-student output with the source DCP student output.
Cosmos-H-Surgical was trained using GraSP real surgical video and synthetic surgical video data. The synthetic collection contains 12,600 laparoscopic cholecystectomy videos generated through the Cosmos-H-Surgical transfer pipeline. GraSP contributes 15,043 robot-assisted radical prostatectomy videos. The combined collection uses a 90% training and 10% validation split.
| Dataset | Source | Role |
|---|---|---|
| Synthetic Data | Synthetic surgical video | Prediction and transfer adaptation |
| GraSP | Real robot-assisted surgery | Surgical-domain adaptation and validation |
The training collection contains approximately 24,879 videos and the held-out validation collection approximately 2,764 videos. Total video duration is less than 10,000 hours.
Other GPU architectures, precisions, and parallelism settings require separate validation.
Please report security vulnerabilities through the NVIDIA Product Security process.