Downloads · 30 days
0
Efficient-Large-Model/LongLive-2.0-5B
LongLive-2.0-5B is a text-to-video model from Efficient-Large-Model. Use it when you need video from a text prompt. It is set up for wan2.2. The card lists the license as other.
<p align="center" <img src="logo.png" alt="LongLive2.0 logo" width="100%" </p
Downloads · 30 days
0
Access
Public
Updated May 19, 2026
Repo size
32.6 GB
Likes
28
Public
Click a slice to open those files.
.pt10 GB · 100%
From the Hugging Face model README
This repository hosts temporary LongLive2.0 5B BF16 checkpoints for inference with the LongLive2.0 release code:
https://github.com/NVlabs/LongLive
The checkpoint package contains two parts:
LongLive2.0 inference loads the base generator first, applies the LoRA modules, and then loads the LoRA weights.
git clone https://github.com/wileewang/LongLive2.0.git
cd LongLive2.0
conda create -n longlive2 python=3.10 -y
conda activate longlive2
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
pip install flash-attn --no-build-isolation
The released LongLive2.0 checkpoint is sufficient for standard inference. You only need to download the original Wan2.2-TI2V-5B components if you want to run training, initialize from the original Wan weights, or use code paths that explicitly load the base Wan model files:
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B \
--local-dir wan_models/Wan2.2-TI2V-5B
Download this checkpoint repository:
huggingface-cli download Perflow-Shuai/longlive_2.0_5B_tmp_20260507 \
--local-dir checkpoints/longlive2_5b
Edit configs/inference.yaml:
checkpoints:
generator_ckpt: checkpoints/longlive2_5b/path/to/base_generator.pt
lora_ckpt: checkpoints/longlive2_5b/path/to/dmd_lora.pt
adapter:
type: lora
rank: 128
alpha: 128
dropout: 0.0
verbose: true
data:
data_path: /path/to/inference_prompts
output_folder: videos/longlive2
num_samples: 1
inference:
sampling_steps: 4
sink_size: 8
guidance_scale: 1.0
multi_shot_sink: true
multi_shot_rope_offset: 8
Replace the checkpoint filenames above with the actual files in this repository.
If the LoRA checkpoint is not used, remove the adapter section and leave
lora_ckpt unset.
data.data_path is passed to MultiTextConcatDataset in inference.py. It can
be either:
.txt file, where each line is one single-shot prompt; orFor a directory input, the code supports both of the following layouts. The direct caption-root layout is the simplest:
inference_prompts/
robot_lab_demo/
0.json
1.json
2.json
shot_durations.txt
It also supports a dataset root with an outer caption/ folder:
inference_prompts/
caption/
robot_lab_demo/
0.json
1.json
2.json
shot_durations.txt
Each JSON file contains:
{
"caption": "A compact silver robot with one blue optic explores a clean robotics lab."
}
shot_durations.txt is optional. If provided, each number is the number of
temporal chunks assigned to the corresponding caption, for example:
2 2 4
Single node, 8 GPUs:
torchrun --standalone --nnodes=1 --nproc_per_node=8 inference.py \
--config_path configs/inference.yaml
Single GPU:
python inference.py --config_path configs/inference.yaml
Outputs are written to output_folder.
inference.sampling_steps controls the number of denoising steps.inference.multi_shot_sink enables the multi-shot attention sink.inference.multi_shot_rope_offset controls the multi-shot RoPE offset.GOVERNING TERMS: This trial service is governed by the NVIDIA API Trial Terms of Service. Use of this model is governed by the NVIDIA Open Model License Agreement.
@article{longlive_2,
title={LongLive2.0: An NVFP4 Parallel Infrastructure for Long Video Generation},
author={Chen, Yukang and Wang, Luozhou and Huang, Wei and Yang, Shuai and Zhang, Bohan and Xiao, Yicheng and Chu, Ruihang and Mao, Weian and Hu, Qixin and Liu, Shaoteng and Zhao, Yuyang and Mao, Huizi and Chen, Ying-Cong and Xie, Enze and Qi, Xiaojuan and Han, Song},
journal={arXiv preprint arXiv},
year={2026}
}