Downloads · 30 days
10
8% of all-time downloads
bytedance-research/EchoVideo
EchoVideo is a machine learning model from bytedance-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as other.
This repo contains PyTorch model definitions, pre-trained weights and inference code for our video generation model, EchoVideo. EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion <be
Downloads · 30 days
10
8% of all-time downloads
All-time downloads
126
Public
Parameters
5.6B
11.3 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors11.2 GB · 99%
From the Hugging Face model README
This repo contains PyTorch model definitions, pre-trained weights and inference code for our video generation model, EchoVideo.
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion <be>
[2025.02.27] We release the inference code and model weights of EchoVideo.
EchoVideo is capable of generating a personalized video from a single photo and a text description. It excels in addressing issues related to "semantic conflict" and "copy-paste" problems. And demonstrates state-of-the-art performance.
| Face-ID Preserving | Full-Body Preserving |
|---|---|
| <img height="300" src="asset/examples/3.gif" > | <img height="300" src="asset/examples/4.gif" > |
| EchoVideo | ConsisID | IDAnimator |
|---|---|---|
| <img height="240" src="asset/examples/2.gif" > | <img height="240" src="asset/examples/5.gif" > | <img height="240" src="asset/examples/6.gif" > |
| <img height="240" src="asset/examples/1.gif" > | <img height="240" src="asset/examples/7.gif" > | <img height="240" src="asset/examples/8.gif" > |
Python version is between 3.10 and 3.12, inclusive of both 3.10 and 3.12. Support both gpu and npu
git clone https://github.com/bytedance/EchoVideo
cd EchoVideo
pip install -r requirements.txt
The details of download pretrained models are shown here.
# multi-resolution video generation [(480, 640), (480, 848), (480, 480), (848, 480), (640, 480)]
python infer.py
Overall architecture of EchoVideo. By employing a meticulously designed IITF module and mitigating the over-reliance on input images, our model effectively unifies the semantic information between the input facial image and the textual prompt. This integration enables the generation of consistent characters with multi-view facial coherence, ensuring that the synthesized outputs maintain both visual and semantic fidelity across diverse perspectives.
Illustration of facial information injection methods. (a) IITF. Facial and textual information are fused to ensure consistent guidance throughout the generation process. we propose IITF to fuse text and facial information, establishing a semantic bridge between facial and textual information, coordinating the influence of different information on character features, thereby ensuring the consistency of generated characters. IITF consists of two core components: facial feature alignment and conditional feature alignment. (b) Dual branch. Facial and textual information are independently injected through Cross Attention mechanisms, providing separate guidance for the generation process.
| Model | Identity Average↑ | Identity Variation↓ | Inception Distance↓ | Dynamic Degree↑ |
|---|---|---|---|---|
| IDAnimator | 0.349 | 0.032 | 159.11 | 0.280 |
| ConsisID | <u>0.414</u> | 0.094 | 200.40 | 0.871 |
| pika | 0.329 | 0.091 | 268.35 | <u>0.954</u> |
| Ours | 0.516 | <u>0.075</u> | <u>176.53</u> | 0.955 |
If you find our work useful in your research, please consider citing the paper
@article{wei2025echovideo,
title={EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion},
author={Wei, Jiangchuan and Yan, Shiyue and Lin, Wenfeng and Liu, Boyuan and Chen, Renjie and Guo, Mingyu},
journal={arXiv preprint arXiv:2501.13452},
year={2025}
}