Downloads · 30 days
0
lixinyizju/moda
moda is a machine learning model from lixinyizju. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<h1 align='center'MoDA: Multi-modal Diffusion Architecture for Talking Head Generation</h1
Downloads · 30 days
0
Access
Public
Updated Aug 8, 2025
Repo size
4.2 GB
Likes
13
Public
Click a slice to open those files.
.pth2.1 GB · 51%
From the Hugging Face model README
<strong>Authors</strong> <br><br>
Xinyang Li<sup>1,2</sup>, Gen Li<sup>2</sup>, Zhihui Lin<sup>1,3</sup>, Yichen Qian<sup>1,3 †</sup>, Gongxin Yao<sup>2</sup>, Weinan Jia<sup>1</sup>, Aowen Wang<sup>1</sup>, Weihua Chen<sup>1,3</sup>, Fan Wang<sup>1,3</sup> <br><br>
<sup>1</sup>Xunguang Team, DAMO Academy, Alibaba Group <sup>2</sup>Zhejiang University <sup>3</sup>Hupan Lab <br><br>
<sup>†</sup>Corresponding authors: [email protected], [email protected]
</div> <br> <div align='center'> <a href='https://lixinyyang.github.io/MoDA.github.io/'><img src='https://img.shields.io/badge/Project-Page-blue'></a> <a href='https://arxiv.org/abs/2507.03256'><img src='https://img.shields.io/badge/Paper-Arxiv-red'></a> </div>Create environment:
# 1. Create base environment
conda create -n moda python=3.10 -y
conda activate moda
# 2. Install requirements
pip install -r requirements.txt
# 3. Install ffmpeg
sudo apt-get update
sudo apt-get install ffmpeg -y
python src/models/inference/moda_test.py --image_path src/examples/reference_images/6.jpg --audio_path src/examples/driving_audios/5.wav
This project is intended for academic research, and we explicitly disclaim any responsibility for user-generated content. Users are solely liable for their actions while using the generative model. The project contributors have no legal affiliation with, nor accountability for, users' behaviors. It is imperative to use the generative model responsibly, adhering to both ethical and legal standards.
We would like to thank the contributors to the LivePortrait, and echomimic,JoyVasa,Ditto, Open Facevid2vid, InsightFace, X-Pose, DiffPoseTalk, Hallo, wav2vec 2.0, Chinese Speech Pretrain, Q-Align, Syncnet, and VBench repositories, for their open research and extraordinary work. If we missed any open-source projects or related articles, we would like to complement the acknowledgement of this specific work immediately.
If you use MoDA in your research, please cite:
@article{li2025moda,
title={MoDA: Multi-modal Diffusion Architecture for Talking Head Generation},
author={Li, Xinyang and Li, Gen and Lin, Zhihui and Qian, Yichen and Yao, GongXin and Jia, Weinan and Chen, Weihua and Wang, Fan},
journal={arXiv preprint arXiv:2507.03256},
year={2025}
}