Downloads · 30 days
9
14% of all-time downloads
malper/hierarcaps-clip-b
hierarcaps-clip-b is a zero-shot image classification model from malper. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Fine-tuned CLIP-B (ViT-B/32) checkpoint from the paper Emergent Visual-Semantic Hierarchies in Image-Text Representations (ECCV 2024).
Downloads · 30 days
9
14% of all-time downloads
All-time downloads
65
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.bin605 MB · 99%
From the Hugging Face model README
Fine-tuned CLIP-B (ViT-B/32) checkpoint from the paper Emergent Visual-Semantic Hierarchies in Image-Text Representations (ECCV 2024).
This model was fine-tuned on the HierarCaps dataset to enhance hierarchical organization in CLIP's embedding space, improving hierarchical reasoning about specificity between images and texts.
from transformers import CLIPModel, AutoProcessor
model = CLIPModel.from_pretrained("malper/hierarcaps-clip-b")
processor = AutoProcessor.from_pretrained("malper/hierarcaps-clip-b")
See the official code repository for training and evaluation scripts used with this checkpoint.
@InProceedings{alper2024hierarcaps,
author = {Morris Alper and Hadar Averbuch-Elor},
title = {Emergent Visual-Semantic Hierarchies in Image-Text Representations},
booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
year = {2024}
}