Downloads · 30 days
72
3% of all-time downloads
GAIR/daVinci-MagiHuman
daVinci-MagiHuman is a image-to-video model from GAIR. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
72
3% of all-time downloads
All-time downloads
2.4K
Public
Repo size
216 GB
Likes
335
Public
Click a slice to open those files.
.safetensors214 GB · 99%
From the Hugging Face model README
This repository contains the weights for daVinci-MagiHuman, introduced in the paper.
<p align="center Visitor"> <a href="https://plms.ai">SII-GAIR</a> & <a href="https://sand.ai">Sand.ai</a> </p> </div>daVinci-MagiHuman uses a single-stream Transformer that takes text tokens, a reference image latent, and noisy video and audio tokens as input, and jointly denoises the video and audio within a unified token sequence.
Key design choices:
| Component | Description |
|---|---|
| Sandwich Architecture | First and last 4 layers use modality-specific projections; middle 32 layers share parameters across modalities |
| Timestep-Free Denoising | No explicit timestep embeddings — the model infers the denoising state directly from input latents |
| Per-Head Gating | Learned scalar gates with sigmoid activation on each attention head for training stability |
| Unified Conditioning | Denoising and reference signals handled through a minimal unified interface — no dedicated conditioning branches |
| Model | Visual Quality ↑ | Text Alignment ↑ | Physical Consistency ↑ | WER ↓ |
|---|---|---|---|---|
| OVI 1.1 | 4.73 | 4.10 | 4.41 | 40.45% |
| LTX 2.3 | 4.76 | 4.12 | 4.56 | 19.23% |
| daVinci-MagiHuman | 4.80 | 4.18 | 4.52 | 14.60% |
| Matchup | daVinci-MagiHuman Win | Tie | Opponent Win |
|---|---|---|---|
| vs Ovi 1.1 | 80.0% | 8.2% | 11.8% |
| vs LTX 2.3 | 60.9% | 17.2% | 21.9% |
| Resolution | Base (s) | Super-Res (s) | Decode (s) | Total (s) |
|---|---|---|---|---|
| 256p | 1.6 | — | 0.4 | 2.0 |
| 540p | 1.6 | 5.1 | 1.3 | 8.0 |
| 1080p | 1.6 | 31.0 | 5.8 | 38.4 |
# Pull the MagiCompiler Docker image
docker pull sandai/magi-compiler:latest
# Launch container
docker run -it --gpus all \
-v /path/to/models:/models \
sandai/magi-compiler:latest bash
# Install MagiCompiler
git clone https://github.com/SandAI-org/MagiCompiler
cd MagiCompiler
pip install -e . --no-build-isolation --config-settings editable_mode=compat
cd ..
# Clone daVinci-MagiHuman
git clone https://github.com/GAIR-NLP/daVinci-MagiHuman
cd daVinci-MagiHuman
# Create environment
conda create -n davinci python=3.12
conda activate davinci
# Install PyTorch
pip install torch==2.9.0 torchvision==0.24.0 torchaudio==2.9.0
# Install Flash Attention (Hopper)
git clone https://github.com/Dao-AILab/flash-attention
cd flash-attention/hopper && python setup.py install && cd ../..
# Install MagiCompiler
git clone https://github.com/SandAI-org/MagiCompiler
cd MagiCompiler
pip install -e . --no-build-isolation --config-settings editable_mode=compat
cd ..
# Clone and install daVinci-MagiHuman
git clone https://github.com/GAIR-NLP/daVinci-MagiHuman
cd daVinci-MagiHuman
pip install -r requirements.txt
Download the complete model stack from HuggingFace and update the paths in the config files under example/.
Before running, update the checkpoint paths in the config files (example/*/config.json) to point to your local model directory.
Base Model (256p)
bash example/base/run.sh
Distilled Model (256p, 8 steps, no CFG)
bash example/distill/run.sh
Super-Resolution to 540p
bash example/sr_540p/run.sh
Super-Resolution to 1080p
bash example/sr_1080p/run.sh
We thank the open-source community, and in particular Wan2.2 and Turbo-VAED, for their valuable contributions.
This project is released under the Apache License 2.0.
@misc{davinci-magihuman-2026,
title = {Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model},
author = {SII-GAIR and Sand.ai},
year = {2026},
url = {https://github.com/GAIR-NLP/daVinci-MagiHuman}
}