Downloads · 30 days
28
2% of all-time downloads
PAIR/text2video-zero-controlnet-canny-avatar
text2video-zero-controlnet-canny-avatar is a text-to-video model from PAIR. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as creativeml-openrail-m.
Text2Video-Zero is a zero-shot text to video generator. It can perform zero-shot text-to-video generation, Video Instruct Pix2Pix (instruction-guided video editing), text and pose conditional video generation, text an…
Downloads · 30 days
28
2% of all-time downloads
All-time downloads
1.8K
Public
Repo size
11 GB
Likes
10
Public
Click a slice to open those files.
.bin5.5 GB · 100%
From the Hugging Face model README
Text2Video-Zero is a zero-shot text to video generator. It can perform zero-shot text-to-video generation, Video Instruct Pix2Pix (instruction-guided video editing),
text and pose conditional video generation, text and canny-edge conditional video generation, and
text, canny-edge and dreambooth conditional video generation. For more information about this work,
please have a look at our paper and our demo:
Our code works with any StableDiffusion base model.
This model provides DreamBooth weights for the Avatar style to be used with edge guidance (using ControlNet) in text2video zero.
We converted the original weights into diffusers and made them usable for ControlNet with edge guidance using: https://github.com/lllyasviel/ControlNet/discussions/12.
Developed by: Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan and Humphrey Shi
Model type: Dreambooth text-to-image and text-to-video generation model with edge control for text2video zero
Language(s): English
License: The CreativeML OpenRAIL M license.
Model Description: This is a model for text2video zero with edge guidance and avatar style. It can be used also with ControlNet in a text-to-image setup with edge guidance.
DreamBoth Keyword: avatar style
Cite as:
@article{text2video-zero,
title={Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators},
author={Khachatryan, Levon and Movsisyan, Andranik and Tadevosyan, Vahram and Henschel, Roberto and Wang, Zhangyang and Navasardyan, Shant and Shi, Humphrey},
journal={arXiv preprint arXiv:2303.13439},
year={2023}
}
The Dreambooth weights for the Avatar style were taken from CIVITAI.
Beware that Text2Video-Zero may output content that reinforces or exacerbates societal biases, as well as realistic faces, pornography, and violence. Text2Video-Zero in this demo is meant only for research purposes.
@article{text2video-zero,
title={Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators},
author={Khachatryan, Levon and Movsisyan, Andranik and Tadevosyan, Vahram and Henschel, Roberto and Wang, Zhangyang and Navasardyan, Shant and Shi, Humphrey},
journal={arXiv preprint arXiv:2303.13439},
year={2023}
}