Downloads · 30 days
5.7K
6% of all-time downloads
NOVAglow646/Monet-7B
Monet-7B is a image-text-to-text model from NOVAglow646. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This is the pretrained model for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language"
Downloads · 30 days
5.7K
6% of all-time downloads
All-time downloads
92.8K
Public
Parameters
8.3B
16.6 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
This is the pretrained model for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language"
Paper: http://arxiv.org/abs/2511.21395
Code: https://github.com/NOVAglow646/Monet
How to use this model: we provide an inference example in our GitHub repo.
If you find this work useful, please use the following BibTeX. Thank you for your support!
@misc{wang2025monetreasoninglatentvisual,
title={Monet: Reasoning in Latent Visual Space Beyond Images and Language},
author={Qixun Wang and Yang Shi and Yifei Wang and Yuanxing Zhang and Pengfei Wan and Kun Gai and Xianghua Ying and Yisen Wang},
year={2025},
eprint={2511.21395},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.21395},
}