Downloads · 30 days
0
glab-caltech/VALOR-GroundingDINO
VALOR-GroundingDINO is a object detection model from glab-caltech. Use it when you need objects located in an image. The card lists the license as mit.
This is the verified-tuned GroundingDINO model from the paper: No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
979 MB
Likes
0
Public
Click a slice to open those files.
.pth979 MB · 100%
From the Hugging Face model README
This is the verified-tuned GroundingDINO model from the paper: No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
For further information please refer to the project webpage, paper, and repository.
If you use VALOR in your research, please consider citing our work:
BibTeX:
@misc{marsili2025labelsproblemtrainingvisual,
title={No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers},
author={Damiano Marsili and Georgia Gkioxari},
year={2025},
eprint={2512.08889},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.08889},
}