Downloads · 30 days
38
69% of all-time downloads
LiconStudio/Licon-Live-Talk
Licon-Live-Talk is a machine learning model from LiconStudio. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Licon Live Talk is a digital human generation model built on the LTX native audio-video generation architecture. Given a reference portrait and text guidance, it generates synchronized audio-video content as a continu…
Downloads · 30 days
38
69% of all-time downloads
All-time downloads
55
Public
Repo size
80.2 GB
Likes
6
Public
Click a slice to open those files.
.safetensors79.8 GB · 100%
From the Hugging Face model README
Licon Live Talk is a digital human generation model built on the LTX native audio-video generation architecture. Given a reference portrait and text guidance, it generates synchronized audio-video content as a continuously extendable stream rather than a fixed-length clip.
The current release demonstrates that this technical direction is feasible. It is intended for research, evaluation, and further development, not production deployment.
The inference runtime, installation steps, and model download instructions are available at:
github.com/liconstudio/Licon-Live-Talk
The generator is adapted with a Self-Forcing training strategy. Instead of relying only on clean teacher trajectories, the model is trained to continue from states produced by its own rollout. This reduces the mismatch between training conditions and real inference, where every new segment must build on content generated by the model itself.
Self-Forcing describes the training strategy. For streaming inference, this release adds an anchor-and-history continuation mechanism organized around two forms of generated context:
Within each segment, new audio-video blocks are generated causally from the anchor, recent history, and current text guidance, while a temporary KV cache advances block by block. When the segment is complete, its latest latent tail becomes the rolling history for the next segment while the identity anchor remains fixed. This process enables the model's infinite-generation mode while keeping the cross-segment context bounded.
Audio and video are generated within the same LTX-based pipeline. They are not produced as two unrelated outputs and combined only at the end, which provides a stronger base for synchronized character performance.
"Infinite generation" refers to a continuously extendable streaming process with no fixed pre-declared clip length. Practical session length still depends on the deployment environment and application.
<video src="https://huggingface.co/LiconStudio/Licon-Live-Talk/resolve/main/validation_0.1.mp4" controls playsinline width="100%"></video>
Open the validation video directly
This validation release focuses on three questions:
The inference code is released under the Apache License 2.0. Bundled third-party components retain their original licenses and terms. See THIRD_PARTY_NOTICES.md in the inference repository.