Downloads · 30 days
2.2K
92% of all-time downloads
akshan-main/tiny-wan22-vace-modular-pipe
tiny-wan22-vace-modular-pipe is a text-to-image model from akshan-main. Use it when you need an image from a text prompt. It is set up for diffusers.
This is a modular diffusion pipeline built with 🧨 Diffusers' modular pipeline framework.
Downloads · 30 days
2.2K
92% of all-time downloads
All-time downloads
2.4K
Public
Parameters
59.4K
851 KB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors851 KB · 67%
From the Hugging Face model README
This is a modular diffusion pipeline built with 🧨 Diffusers' modular pipeline framework.
Pipeline Type: Wan22VaceBlocks
Description: Modular pipeline for controllable video generation using Wan2.2 VACE.
This pipeline uses a 5-block architecture that can be customized and extended.
[TODO]
This modular pipeline is composed of the following blocks:
WanTextEncoderStep)
WanVaceEncoderStep)
Wan22VaceCoreDenoiseStep)
WanVaceTrimReferenceLatentsStep)
WanVaeDecoderStep)
UMT5EncoderModel)AutoTokenizer)ClassifierFreeGuidance)WanVACETransformer3DModel)AutoencoderKLWan)VideoProcessor)UniPCMultistepScheduler)ClassifierFreeGuidance)WanVACETransformer3DModel)boundary_ratio (default: 0.875): The boundary ratio to divide the denoising loop into high noise and low noise stages.
Inputs:
prompt (None, optional): No description providednegative_prompt (None, optional): No description providedmax_sequence_length (None, optional, defaults to 512): No description providedvideo (list, optional): The control video to condition the generation on. If not provided, an empty video is used.mask (list, optional): The mask that defines which video regions to condition on (black) and which to generate (white). Can only be passed if video is passed as well.reference_images (Image | list, optional): One or more reference images as extra conditioning for the generation.conditioning_scale (float | list | Tensor, optional, defaults to 1.0): The conditioning scale applied in each control layer of the model. If a float, it is applied uniformly to all layers; a list or tensor must have the same length as the number of control layers.height (None, optional): No description providedwidth (None, optional): No description providednum_frames (int, optional, defaults to 81): No description providedgenerator (None, optional): No description providednum_videos_per_prompt (None, optional, defaults to 1): No description providednum_inference_steps (None, optional, defaults to 50): No description providedtimesteps (None, optional): No description providedsigmas (None, optional): No description providedlatents (Tensor | NoneType, optional): No description providedattention_kwargs (None, optional): No description providedoutput_type (str, optional, defaults to np): The output type of the decoded videosOutputs:
videos (list): The generated videos.