Downloads · 30 days
2
4% of all-time downloads
BornFly/AsymmetricMagVitV2_4z
AsymmetricMagVitV2_4z is a machine learning model from BornFly. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Downloads · 30 days
2
4% of all-time downloads
All-time downloads
55
Public
Repo size
1.4 GB
Likes
2
Public
Click a slice to open those files.
.bin517 MB · 91%
From the Hugging Face model README
Lightweight open-source reproduction of MagVitV2, fully aligned with the paper’s functionality. Supports image and video joint encoding and decoding, as well as videos of arbitrary length and resolution.
github_link: https://github.com/bornfly-detachment/asymmetric_magvitv2
Original video (above) & VAE Reconstruction video (below)
<table> <tr> <td width="50%"> <a href="https://github.com/bornfly-detachment/asymmetric_magvitv2/assets/174133722/2ec0bc1b-7a32-4949-a68f-d512bd4c5411"> <img src="https://github.com/bornfly-detachment/asymmetric_magvitv2/assets/174133722/2ec0bc1b-7a32-4949-a68f-d512bd4c5411" alt="60s 3840x2160" style="width: 100%;"> </a> </td> <td width="50%"> <a href="https://github.com/bornfly-detachment/asymmetric_magvitv2/assets/174133722/c8df652d-e7b9-42ff-a554-1271d5a0bce1"> <img src="https://github.com/bornfly-detachment/asymmetric_magvitv2/assets/174133722/c8df652d-e7b9-42ff-a554-1271d5a0bce1" alt="60s 1920x1080" style="width: 100%;"> </a> </td> </tr> </table>bilibili_Black Myth:Wu KongULR 4zVAE
<a name="installation"></a>
git clonehttps://github.com/bornfly-detachment/AsymmetricMagVitV2.git
cd AsymmetricMagVitV2
This is assuming you have navigated to the AsymmetricMagVitV2 root after cloning it.
# install required packages from pypi
python3 -m venv .pt2
source .pt2/bin/activate
pip3 install -r requirements/pt2.txt
| model | downsample (THW) | Encoder Size | Decoder Size |
|---|---|---|---|
| svd 2Dvae | 1x8x8 | 34M | 64M |
| AsymmetricMagVitV2 | 4x8x8 | 100M | 159M |
| model | Data | #iterations | URL |
|---|---|---|---|
| AsymmetricMagVitV2_4z | 20M Intervid | 2node 1200k | AsymmetricMagVitV2_4z |
| AsymmetricMagVitV2_16z | 20M Intervid | 4node 860k | AsymmetricMagVitV2_16z |
<a name="Metric"></a>
| model | temporal-frame | fvd(↓) | fid(↓) | psnr(↑) | ssim(↑) |
|---|---|---|---|---|---|
| SVD VAE | 1 | 190.6 | 1.8 | 28.2 | 1.0 |
| openSoraPlan | 1 | 249.8 | 1.04 | 29.6 | 0.99 |
| openSoraPlan | 17 | 725.4 | 3.17 | 23.4 | 0.89 |
| openSoraPlan | 33 | 756.8 | 3.5 | 23 | 0.89 |
| AsymmetricMagVitV2_4z | 1 | 113.5 | 1.4 | 29.8 | 1.0 |
| AsymmetricMagVitV2_4z | 17 | 278.5 | 2.3 | 26.4 | 1.0 |
| AsymmetricMagVitV2_4z | 33 | 293.3 | 2.5 | 26.3 | 1.0 |
Note:
If the GPU VRAM is not sufficient, metrics for evaluation can be adjusted to be between 256 and 512 at maximum.
(default GPU VRAM needs to exceed 28GB. If the GPU VRAM is not sufficient, metrics for evaluation can be adjusted to be between 32=256p/8 and 64=512p/8 at maximum.)
5 frames of latent space corresponds to 17 frames of the original video. The calculation formula is as follows: latent_T_dim = (frame_T_dim - 1) / temporal_downsample_num; in this model, temporal_downsample_num=4
from models.vae import AsymmetricMagVitV2Pipline
import torch
from models.utils.image_op import imdenormalize, imnormalize, read_video, read_image
import torchvision.transforms as transforms
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16
encoder_init_window = 17
input_path = "data/videos/tokyo_walk.mp4"
img_transform = transforms.Compose([transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))])
input, last_frame_id = read_video(input_path, encoder_init_window, sample_fps=8, img_transform, start=0)
model = AsymmetricMagVitV2Pipline.from_pretrained("BornFly/AsymmetricMagVitV2_4z").to(device, dtype).eval()
init_z, reg_log = model.encode(input, encoder_init_window, is_init_image=True, return_reg_log=True, unregularized=False)
init_samples = model.decode(init_z.to(device, dtype), decode_batch_size=1, is_init_image=True)
python infer_vae.py --input_path data/videos/tokyo-walk.mp4 --model_path vae_16z_bf16_hf --output_folder vae_eval_out/vae_4z_bf16_hf_videos > infer_vae_video.log 2>&1
python infer_vae.py --input_path data/images --model_path vae_16z_bf16_hf --output_folder vae_eval_out/vae_4z_bf16_hf_images > infer_vae_image.log 2>&1
Reproducing Sora, a 16-channel VAE integrated with SD3. Due to limited computational resources, the focus is on generating 1K high-definition dynamic wallpapers.
Reproducing VideoPoet, supporting multimodal joint representation. Due to limited computational resources, the focus is on generating music videos.