Downloads · 30 days
37
28% of all-time downloads
ai9stars/Cheers_edit
Cheers_edit is a machine learning model from ai9stars. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
37
28% of all-time downloads
All-time downloads
134
Public
Parameters
2.8B
5.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5.6 GB · 100%
From the Hugging Face model README
Yichen Zhang<sup>1*</sup>, Da Peng<sup>2*</sup>, Zonghao Guo<sup>1†</sup>, Zijian Zhang<sup>3</sup>, Xuesong Yang<sup>3</sup>,
Tong Sun<sup>3</sup>, Shichu Sun<sup>3</sup>, Yidan Zhang<sup>3</sup>, Yanghao Li<sup>1</sup>, Haiyan Zhao<sup>1</sup>, Wang Xu<sup>1</sup>,
Qi Shi<sup>1</sup>, Yangang Sun<sup>1</sup>, Chi Chen<sup>1</sup>, Shuo Wang<sup>1</sup>, Yukun Yan<sup>1</sup>, Xu Han<sup>1</sup>,
Qiang Ma<sup>1</sup>, Wei Ke<sup>2</sup>, Liang Wang<sup>3</sup>, Zhiyuan Liu<sup>1</sup>, Maosong Sun<sup>1</sup>
<sup>1</sup>Tsinghua University, <sup>2</sup>Xi'an Jiaotong University, <sup>3</sup>University of Chinese Academy of Sciences
* Equal contribution † Corresponding author
A recent cutting-edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes and visual representations, making it non-trivial to jointly optimize within a shared feature space. In this work, we present Cheers, a unified multimodal model that decouples patch-level details from semantic representations, thereby stabilizing semantics for multimodal understanding and improving fidelity for image generation via gated detail residuals. Cheers includes three key components: (i) a unified vision tokenizer that encodes and compresses image latent states into semantic tokens for efficient LLM conditioning, (ii) an LLM-based Transformer that unifies autoregressive decoding for text generation and diffusion decoding for image generation, and (iii) a cascaded flow matching head that decodes visual semantics first and then injects semantically gated detail residuals from the vision tokenizer to refine high-frequency content. Experiments on popular benchmarks demonstrate that Cheers matches or surpasses advanced UMMs in both visual understanding and generation. Notably, Cheers outperforms the Tar-1.5B on the popular benchmarks GenEval and MMBench, while requiring only 20% of the training cost, indicating effective and efficient (i.e., 4x token compression) unified multimodal modeling.
<!-- Provide the basic links for the model. -->
For any questions or collaborations, feel free to contact us : )
<p align="left"> 📧 <a href="[email protected]">[email protected]</a>   |    📧 <a href="[email protected]">[email protected]</a>   |    📧 <a href="[email protected]">[email protected]</a>   </p>If you find Cheers useful, please cite Cheers technical report using this BibTeX.
@article{zhang2026cheers,
title={CHEERS: DECOUPLING PATCH DETAILS FROM SEMANTIC REPRESENTATIONS ENABLES UNIFIED MULTIMODAL COMPREHENSION AND GENERATION},
author={Zhang, Yichen and Peng, Da and Guo, Zonghao and Zhang, Zijian and Yang, Xuesong and Sun, Tong and Sun, Shichu and Zhang, Yidan and Li, Yanghao and Zhao, Haiyan and others},
journal={arXiv preprint arXiv:2603.12793},
year={2026}
}