Downloads · 30 days
0
mkjia/MGVQ
MGVQ is a image-to-image model from mkjia. Use it when you need one image transformed into another.
Downloads · 30 days
0
Access
Public
Updated Sep 20, 2025
Repo size
18.2 GB
Likes
2
Public
Click a slice to open those files.
.npz9.8 GB · 54%
From the Hugging Face model README
Mingkai Jia<sup>1,2</sup>, Wei Yin<sup>2*§</sup>, Xiaotao Hu<sup>1,2</sup>, Jiaxin Guo<sup>3</sup>, Xiaoyang Guo<sup>2</sup><br> Qian Zhang<sup>2</sup>, Xiao-Xiao Long<sup>4</sup>, Ping Tan<sup>1</sup><br>
HKUST<sup>1</sup>, Horizon Robotics<sup>2</sup>, CUHK<sup>3</sup>, NJU<sup>4</sup><br> <sup>*</sup> Corresponding Author, <sup>§</sup> Project Leader <br><br><image src='https://huggingface.co/mkjia/MGVQ/resolve/main/assets/teaser.png'/>
</div>[August 2025] Achieve SOTA at paperwithcode leaderboards: Image Reconstruction on ImageNet and UHDBench. <image src='https://huggingface.co/mkjia/MGVQ/raw/main/assets/SOTA_recon_fid_imagenet_badge.jpg'/> <image src='https://huggingface.co/mkjia/MGVQ/raw/main/assets/SOTA_recon_PSNR_UHD_badge.jpg'/>[August 2025] Released Inference Code[August 2025] Released model zoo.[August 2025] Released dataset for ultra-high-definition image reconstruction evaluation. Our proposed super-resolution image reconstruction UHDBench dataset is released.[July 2025] Released paper.| Model | Downsample | Groups | Codebook Size | Training Data | Link |
|---|---|---|---|---|---|
| mgvq-f8c32-g4 | 8 | 4 | 32768 | imagenet | link |
| mgvq-f8c32-g8 | 8 | 8 | 16384 | imagenet | link |
| mgvq-f16c32-g4 | 16 | 4 | 32768 | imagenet | link |
| mgvq-f16c32-g8 | 16 | 8 | 16384 | imagenet | link |
| mgvq-f16c32-g4-mix | 16 | 4 | 32768 | mix | link |
| mgvq-f32c32-g8-mix | 32 | 8 | 16384 | mix | link |
<a id="quick start"></a>
git clone https://github.com/MKJia/MGVQ.git
cd MGVQ
pip3 install requirements.txt
Download the pretrained models from our model zoo to your /path/to/your/ckpt.
Try our UHDBench dataset on huggingface and download to your /path/to/your/dataset.
Remember to change the paths of ckpt and dataset_root, and make sure you are evaluating the expected model on dataset.
cd evaluation
python3 eval_recon.sh
You can download the pretrained GPT model for generation on huggingface, and test it with our mgvq-f16c32-g4 tokenizer model for demo image sampling. Remember to change the paths of gpt_ckpt and vq_ckpt.
cd evaluation
python3 demo_gen.sh
We also provide our .npz file on huggingface sampled by sample_c2i_ddp.py for evaluation.
cd evaluation
python3 evaluator.py /path/to/your/VIRTUAL_imagenet256_labeled.npz /path/to/your/GPT_XXL_300ep_topk_12.npz
If the paper and code from MGVQ help your research, we kindly ask you to give a citation to our paper ❤️. Additionally, if you appreciate our work and find this repository useful, giving it a star ⭐️ would be a wonderful way to support our work. Thank you very much.
@article{jia2025mgvq,
title={MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization},
author={Jia, Mingkai and Yin, Wei and Hu, Xiaotao and Guo, Jiaxin and Guo, Xiaoyang and Zhang, Qian and Long, Xiao-Xiao and Tan, Ping},
journal={arXiv preprint arXiv:2507.07997},
year={2025}
}
This repository is under the MIT License. For more license questions, please contact Mingkai Jia ([email protected]) and Wei Yin ([email protected]).