Downloads · 30 days
0
marcosremar2/MuseTalk1.5
MuseTalk1.5 is a machine learning model from marcosremar2. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers.
<strongMuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling</strong
Downloads · 30 days
0
Access
Public
Updated Dec 17, 2025
Repo size
8.1 GB
Likes
0
Public
Click a slice to open those files.
.pth3.9 GB · 47%
From the Hugging Face model README
<strong>MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling</strong>
Yue Zhang<sup>*</sup>, Zhizhou Zhong<sup>*</sup>, Minhao Liu<sup>*</sup>, Zhaokang Chen, Bin Wu<sup>†</sup>, Yubin Zeng, Chao Zhan, Junxin Huang, Yingjie He, Wenjiang Zhou (<sup>*</sup>Equal Contribution, <sup>†</sup>Corresponding Author, [email protected])
Lyra Lab, Tencent Music Entertainment
github huggingface space Technical report
We introduce MuseTalk, a real-time high quality lip-syncing model (30fps+ on an NVIDIA Tesla V100). MuseTalk can be applied with input videos, e.g., generated by MuseV, as a complete virtual human solution.
We're excited to unveil MuseTalk 1.5. This version (1) integrates training with perceptual loss, GAN loss, and sync loss, significantly boosting its overall performance. (2) We've implemented a two-stage training strategy and a spatio-temporal data sampling approach to strike a balance between visual quality and lip-sync accuracy. Learn more details here
MuseTalk is a real-time high quality audio-driven lip-syncing model trained in the latent space of ft-mse-vae, which
256 x 256.
MuseTalk was trained in latent spaces, where the images were encoded by a freezed VAE. The audio was encoded by a freezed
whisper-tiny model. The architecture of the generation network was borrowed from the UNet of the stable-diffusion-v1-4, where the audio embeddings were fused to the image embeddings by cross-attention.
Note that although we use a very similar architecture as Stable Diffusion, MuseTalk is distinct in that it is NOT a diffusion model. Instead, MuseTalk operates by inpainting in the latent space with a single step.
https://github.com/TMElyralab/MuseTalk/assets/163980830/37a3a666-7b90-4244-8d3a-058cb0e44107
https://github.com/user-attachments/assets/1ce3e850-90ac-4a31-a45f-8dfa4f2960ac
https://github.com/user-attachments/assets/fa3b13a1-ae26-4d1d-899e-87435f8d22b3
https://github.com/user-attachments/assets/15800692-39d1-4f4c-99f2-aef044dc3251
https://github.com/user-attachments/assets/a843f9c9-136d-4ed4-9303-4a7269787a60
https://github.com/user-attachments/assets/6eb4e70e-9e19-48e9-85a9-bbfa589c5fcb
</td> <td width="33%">https://github.com/user-attachments/assets/c04f3cd5-9f77-40e9-aafd-61978380d0ef
https://github.com/user-attachments/assets/2051a388-1cef-4c1d-b2a2-3c1ceee5dc99
https://github.com/user-attachments/assets/b5f56f71-5cdc-4e2e-a519-454242000d32
https://github.com/user-attachments/assets/a5843835-04ab-4c31-989f-0995cfc22f34
https://github.com/user-attachments/assets/3dc7f1d7-8747-4733-bbdd-97874af0c028
https://github.com/user-attachments/assets/3c78064e-faad-4637-83ae-28452a22b09a
</td> <td width="33%">https://github.com/user-attachments/assets/999a6f5b-61dd-48e1-b902-bb3f9cbc7247
https://github.com/user-attachments/assets/d26a5c9a-003c-489d-a043-c9a331456e75
https://github.com/user-attachments/assets/471290d7-b157-4cf6-8a6d-7e899afa302c
https://github.com/user-attachments/assets/1ee77c4c-8c70-4add-b6db-583a12faa7dc
https://github.com/user-attachments/assets/370510ea-624c-43b7-bbb0-ab5333e0fcc4
https://github.com/user-attachments/assets/b011ece9-a332-4bc1-b8b7-ef6e383d7bde
</td> </tr> </table>We provide an automated benchmark test for measuring end-to-end latency in real-time avatar conversations. The benchmark measures:
| Component | Description |
|---|---|
| STT | Speech-to-text transcription time (Whisper) |
| LLM | Language model response time |
| TTS | Text-to-speech synthesis time |
| MuseTalk | Video generation time (first frame + total) |
# Start the server
bash server/start_fast.sh
# Run the benchmark (requires Playwright)
python tests/demotalk_benchmark.py "http://localhost:3000/?demo=true"
BREAKDOWN DA LATENCIA (tempos individuais):
1. STT (Whisper): 439 ms
2. LLM (Groq): 247 ms
3. TTS (ElevenLabs): 438 ms
4. MuseTalk 1 Frame: 378 ms
---------------------------------
Latencia Resposta: 1502 ms
MUSETALK (geracao de video):
1 Frame: 378 ms
Total: 5755 ms
Frames gerados: 102
Velocidade: 17.7 fps
Avaliacao: B (aceitavel)
Reference: Human conversation latency is ~250ms.
MuseTalk now includes automatic optimizations for NVIDIA Blackwell GPUs (RTX 50 series) that provide significant speedups with zero quality degradation.
| Mode | Batch=1 | Batch=8 | Speedup vs FP32 |
|---|---|---|---|
| FP32 | 19.56 ms (51 FPS) | 39.04 ms (205 FPS) | baseline |
| FP16 | 19.61 ms (51 FPS) | 20.28 ms (395 FPS) | 1.93x |
| Optimized | 4.75 ms (210 FPS) | 13.94 ms (574 FPS) | 2.8-4.1x |
| Metric | Value | Assessment |
|---|---|---|
| PSNR | 59.85 dB | Excellent (>40 dB = virtually identical) |
| SSIM | 1.0000 | Perfect |
| Rating | A | Zero quality loss |
The following optimizations are automatically enabled on Blackwell GPUs:
| Optimization | Description | Impact |
|---|---|---|
| FP16 Precision | Half-precision for VAE and UNet | 1.5-2x speedup, 27% less VRAM |
| TF32 MatMuls | TensorFloat-32 for matrix operations | ~1.2x speedup |
| torch.compile | JIT compilation with max-autotune | 1.3-2x additional speedup |
Optimizations are enabled by default when running on RTX 5090:
# Automatic - just use the engine normally
from server.fast_engine import initialize_engine
engine = initialize_engine() # Blackwell optimizations auto-applied
To disable (for debugging or comparison):
engine = FastMuseTalkEngine()
engine.load_models(use_blackwell_optimizations=False)
To verify performance on your hardware:
python tests/benchmark_mxfp4.py
For detailed documentation, see docs/MXFP4_OPTIMIZATION.md.
We provide a detailed tutorial about the installation and the basic usage of MuseTalk for new users:
Thanks for the third-party integration, which makes installation and use more convenient for everyone. We also hope you note that we have not verified, maintained, or updated third-party. Please refer to this project for specific results.
To prepare the Python environment and install additional packages such as opencv, diffusers, mmcv, etc., please follow the steps below:
We recommend a python version >=3.10 and cuda version =11.7. Then build environment as follows:
pip install -r requirements.txt
pip install --no-cache-dir -U openmim
mim install mmengine
mim install "mmcv>=2.0.1"
mim install "mmdet>=3.1.0"
mim install "mmpose>=1.1.0"
Download the ffmpeg-static and
export FFMPEG_PATH=/path/to/ffmpeg
for example:
export FFMPEG_PATH=/musetalk/ffmpeg-4.4-amd64-static
You can download weights manually as follows:
# !pip install -U "huggingface_hub[cli]"
export HF_ENDPOINT=https://hf-mirror.com
huggingface-cli download TMElyralab/MuseTalk --local-dir models/
Finally, these weights should be organized in models as follows:
./models/
├── musetalk
│ └── musetalk.json
│ └── pytorch_model.bin
├── musetalkV15
│ └── musetalk.json
│ └── unet.pth
├── dwpose
│ └── dw-ll_ucoco_384.pth
├── face-parse-bisent
│ ├── 79999_iter.pth
│ └── resnet18-5c106cde.pth
├── sd-vae-ft-mse
│ ├── config.json
│ └── diffusion_pytorch_model.bin
└── whisper
├── config.json
├── pytorch_model.bin
└── preprocessor_config.json
We provide inference scripts for both versions of MuseTalk:
sh inference.sh v1.5
This inference script supports both MuseTalk 1.5 and 1.0 models:
configs/inference/test.yaml is the path to the inference configuration file, including video_path and audio_path. The video_path should be either a video file, an image file or a directory of images.
sh inference.sh v1.0
You are recommended to input video with 25fps, the same fps used when training the model. If your video is far less than 25fps, you are recommended to apply frame interpolation or directly convert the video to 25fps using ffmpeg.
:mag_right: We have found that upper-bound of the mask has an important impact on mouth openness. Thus, to control the mask region, we suggest using the bbox_shift parameter. Positive values (moving towards the lower half) increase mouth openness, while negative values (moving towards the upper half) decrease mouth openness.
You can start by running with the default configuration to obtain the adjustable value range, and then re-run the script within this range.
For example, in the case of Xinying Sun, after running the default configuration, it shows that the adjustable value rage is [-9, 9]. Then, to decrease the mouth openness, we set the value to be -7.
python -m scripts.inference --inference_config configs/inference/test.yaml --bbox_shift -7
:pushpin: More technical details can be found in bbox_shift.
</details>As a complete solution to virtual human generation, you are suggested to first apply MuseV to generate a video (text-to-video, image-to-video or pose-to-video) by referring this. Frame interpolation is suggested to increase frame rate. Then, you can use MuseTalk to generate a lip-sync video by referring this.
python -m scripts.realtime_inference --inference_config configs/inference/realtime.yaml --batch_size 4
configs/inference/realtime.yaml is the path to the real-time inference configuration file, including preparation, video_path , bbox_shift and audio_clips.
preparation to True in realtime.yaml to prepare the materials for a new avatar. (If the bbox_shift has changed, you also need to re-prepare the materials.)avatar will use an audio clip selected from audio_clips to generate video.
Inferring using: data/audio/yongen.wav
preparation to False and run this script if you want to genrate more videos using the same avatar.python -m scripts.realtime_inference --inference_config configs/inference/realtime.yaml --skip_save_images
</details>
Thanks for open-sourcing!
Resolution: Though MuseTalk uses a face region size of 256 x 256, which make it better than other open-source methods, it has not yet reached the theoretical resolution bound. We will continue to deal with this problem.
If you need higher resolution, you could apply super resolution models such as GFPGAN in combination with MuseTalk.
Identity preservation: Some details of the original face are not well preserved, such as mustache, lip shape and color.
Jitter: There exists some jitter as the current pipeline adopts single-frame generation.
@article{musetalk,
title={MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling},
author={Zhang, Yue and Zhong, Zhizhou and Liu, Minhao and Chen, Zhaokang and Wu, Bin and Zeng, Yubin and Zhan, Chao and He, Yingjie and Huang, Junxin and Zhou, Wenjiang},
journal={arxiv},
year={2025}
}
code: The code of MuseTalk is released under the MIT License. There is no limitation for both academic and commercial usage.model: The trained model are available for any purpose, even commercially.other opensource model: Other open-source models used must comply with their license, such as whisper, ft-mse-vae, dwpose, S3FD, etc..AIGC: This project strives to impact the domain of AI-driven video generation positively. Users are granted the freedom to create videos using this tool, but they are expected to comply with local laws and utilize it responsibly. The developers do not assume any responsibility for potential misuse by users.