Downloads · 30 days
18
11% of all-time downloads
leafyseay/LaME-2B
LaME-2B is a feature extraction model from leafyseay. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
LaME (Learning to Think in Latent Space for Multimodal Embedding) model based on Qwen2-VL-2B-Instruct.
Downloads · 30 days
18
11% of all-time downloads
All-time downloads
160
Public
Parameters
2.8B
6.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.4 GB · 100%
From the Hugging Face model README
LaME (Learning to Think in Latent Space for Multimodal Embedding) model based on Qwen2-VL-2B-Instruct.
LaME augments Qwen-VL with learnable [REASON] tokens and a latent decoder supervision, jointly optimizing generation and embedding through an information bottleneck. It produces both discriminative and generative multimodal embeddings for text, images, videos, and visual documents.
Without bells and whistles, LaME achieves state-of-the-art multimodal retrieval performance on MMEB-v2 (image / video / visual-document / full aggregate) and MRMR.
[REASON] tokensSee the LaME repository for inference and evaluation examples.
from transformers import AutoModel, AutoProcessor
model = AutoModel.from_pretrained("leafyseay/LaME-2B", trust_remote_code=True, torch_dtype="bfloat16").cuda()
processor = AutoProcessor.from_pretrained("leafyseay/LaME-2B", trust_remote_code=True)
@article{wu2026lame,
title = {LaME: Learning to Think in Latent Space for Multimodal Embedding via Information Bottleneck},
author = {Wu, Peixi and Yang, Biao and Ma, Feipeng and Chai, Bosong and Lin, Bo and Yuan, Wei and Yang, Fan and Gao, Tingting and Li, Hebei and Sun, Xiaoyan},
journal = {arXiv preprint arXiv:2606.13061},
year = {2026}
}