Downloads · 30 days
169
10% of all-time downloads
fushh7/ObjEmbed-2B
ObjEmbed-2B is a image feature extraction model from fushh7. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
ObjEmbed is a multimodal embedding model designed to align specific image regions (objects) with textual descriptions. Unlike global embedding models, ObjEmbed decomposes an image into multiple regional embeddings alo…
Downloads · 30 days
169
10% of all-time downloads
All-time downloads
1.7K
Public
Parameters
2.4B
5.4 GB on disk
Likes
1
Trending 1
Click a slice to open those files.
.safetensors5.4 GB · 100%
From the Hugging Face model README
ObjEmbed is a multimodal embedding model designed to align specific image regions (objects) with textual descriptions. Unlike global embedding models, ObjEmbed decomposes an image into multiple regional embeddings along with global embeddings, supporting tasks such as visual grounding, local image retrieval, and global image retrieval.
If you find ObjEmbed helpful for your research, please consider citing:
@article{fu2026objembed,
title={ObjEmbed: Towards Universal Multimodal Object Embeddings},
author={Fu, Shenghao and Su, Yukun and Rao, Fengyun and LYU, Jing and Xie, Xiaohua and Zheng, Wei-Shi},
journal={arXiv preprint arXiv:2602.01753},
year={2026}
}