Downloads · 30 days
20
100% of all-time downloads
CurioSeaLab/SoulX-FlashHead-1_3B
SoulX-FlashHead-1_3B is a image-to-video model from CurioSeaLab. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
Soul-AILab/SoulX-FlashHead-13B <div align="center"
Downloads · 30 days
20
100% of all-time downloads
All-time downloads
20
Public
Repo size
14.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors13.8 GB · 96%
From the Hugging Face model README
Soul-AILab/SoulX-FlashHead-1_3B
<div align="center"> <h1>SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads</h1>Tan Yu*, Qian Qiao*<sup>✉</sup>, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu <sup>✉</sup>
<sup>*</sup>Equal Contribution <sup>✉</sup>Corresponding Author
<a href='https://soul-ailab.github.io/soulx-flashhead/' target="_blank"><img src='https://img.shields.io/badge/Project-Page-green'></a> <a href='https://arxiv.org/pdf/2602.07449' target="_blank"><img src='https://img.shields.io/badge/Technical-Report-red'></a> <a href='https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B' target="_blank"><img src='https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-blue'></a> <a href="https://huggingface.co/datasets/Soul-AILab/VividHead" target="_blank"><img src="https://img.shields.io/badge/🤗 Hugging Face-Dataset-blue" alt="Dataset"></a>
</div>More examples are available in the project.
<table> <tbody> <!-- Row 1: Videos 1-5 --> <tr> <td width="30%"><video src="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B/resolve/main/assets/qitiandasheng.mp4" style="width:100%; aspect-ratio:512/512; object-fit:cover;" controls loop></video></td> <td width="30%"><video src="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B/resolve/main/assets/chengdu.mp4" style="width:100%; aspect-ratio:512/512; object-fit:cover;" controls loop></video></td> <td width="30%"><video src="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B/resolve/main/assets/einstein.mp4" style="width:100%; aspect-ratio:512/512; object-fit:cover;" controls loop></video></td> </tr> </tbody> </table>conda create -n flashhead python=3.10
conda activate flashhead
pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
pip install ninja
pip install flash_attn==2.8.0.post2 --no-build-isolation
-- If it takes a long time, we recommend the way below.
pip install sageattention==2.2.0 --no-build-isolation
# Ubuntu / Debian
apt-get install ffmpeg
# CentOS / RHEL
yum install ffmpeg ffmpeg-devel
or
# Conda (no root required)
conda install -c conda-forge ffmpeg==7
# If you are in china mainland, run this first: export HF_ENDPOINT=https://hf-mirror.com
pip install "huggingface_hub[cli]"
huggingface-cli download Soul-AILab/SoulX-FlashHead-1_3B --local-dir ./models/SoulX-FlashHead-1_3B
huggingface-cli download facebook/wav2vec2-base-960h --local-dir ./models/wav2vec2-base-960h
# Infer with [Pro-Model] on single GPU
bash inference_script_single_gpu_pro.sh
# Infer with [Pro-Model] on multy GPUs
bash inference_script_multi_gpu_pro.sh
# Real-time inference speed of Pro-Model can only be supported on two RTX-5090 with SageAttention.
# Infer with [Lite-Model] on single GPU
bash inference_script_single_gpu_lite.sh
# Real-time inference speed can be supported on single RTX-4090 (up to 3 concurrent).
For a real-time interactive experience, scan the QR code to enter the event link. [2026.2.12~2026.3.11] <a id="online-experience-qr"></a>
<div align="center"> <table> <tr> <td align="center"> <img src="assets/soul_event_link.png" width="200" alt="SoulApp event QR Code"/> <br /> <strong>Real-time Online Experience<br>(SoulApp 实时在线体验)</strong> </td> </tr> </table> </div>If you are interested in leaving a message to our work, feel free to email [email protected] or [email protected] or [email protected] or [email protected] or [email protected]
We have opened a WeChat group. Additionally, we represent SoulApp and warmly welcome everyone to download the app and join our Soul group for further technical discussions and updates!
<div align="center"> <table> <tr> <td align="center"> <img src="assets/wechat_group.png" width="300" alt="WeChat Group QR Code"/> <br /> <strong>Join WeChat Group<br>(加入微信技术群)</strong> </td> <td width="100"></td> <td align="center"> <img src="assets/soul_group.png" width="300" alt="Soul App Group QR Code"/> <br /> <strong>Download SoulApp & Join Group<br>(下载SoulApp加入群组)</strong> </td> </tr> </table> </div>If you find our work useful in your research, please consider citing:
@misc{yu2026soulxflashheadoracleguidedgenerationinfinite,
title={SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads},
author={Tan Yu and Qian Qiao and Le Shen and Ke Zhou and Jincheng Hu and Dian Sheng and Bo Hu and Haoming Qin and Jun Gao and Changhai Zhou and Shunshun Yin and Siyuan Liu},
year={2026},
eprint={2602.07449},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.07449},
}
[!TIP] If you find our work useful, please also consider starring the original repositories of these foundational methods.