Downloads · 30 days
0
HaominPeng/EOVSAM
EOVSAM is a image segmentation model from HaominPeng. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. The card lists the license as other.
EOVSAM is an efficient open-vocabulary segmentation framework built on SAM 3 that adapts SAM 3 for single-pass prediction.
Downloads · 30 days
0
Access
Public
Updated Aug 8, 2026
Repo size
7.7 GB
Likes
0
Public
Click a slice to open those files.
.pth7.7 GB · 100%
From the Hugging Face model README
EOVSAM is an efficient open-vocabulary segmentation framework built on SAM 3 that adapts SAM 3 for single-pass prediction.
EOVSAM removes prompt conditioning to turn SAM 3 into an efficient mask generator and introduces an Attentional Aggregation strategy to optimize open-vocabulary classification end-to-end. This formulation avoids multi-stage pipelines and post-processing heuristics while consistently improving segmentation accuracy over vanilla SAM 3 and accelerating inference by up to 338×.
Please refer to the official EOVSAM GitHub repository for installation, dataset preparation, training, and evaluation scripts.
This project is a composite distribution incorporating NVIDIA RADIO, SAM 3, and MAFT-Plus components; each remains subject to its upstream terms. Please check the GitHub License section for details.
@misc{peng2026eovsamefficientopenvocabularysegmentation,
title={EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass},
author={Haomin Peng and Yongkang Li and Zhaoxiang Liu and Xiaojie Jin and Shiguo Lian and Yunchao Wei and Xinggang Wang},
year={2026},
eprint={2608.02284},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.02284},
}