Downloads · 30 days
96
12% of all-time downloads
MVRL/rcme-tol-vit-base-patch16
rcme-tol-vit-base-patch16 is a zero-shot image classification model from MVRL. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for open_clip. The card lists the license as mit.
This model was presented in the paper Global and Local Entailment Learning for Natural World Imagery.
Downloads · 30 days
96
12% of all-time downloads
All-time downloads
773
Public
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.bin599 MB · 50%
From the Hugging Face model README
This model was presented in the paper Global and Local Entailment Learning for Natural World Imagery.
Learning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at this https URL .
Find more information and visualizations on the project page.