Downloads · 30 days
4
2% of all-time downloads
cs-giung/clip-vit-base-patch32-fullcc2.5b
clip-vit-base-patch32-fullcc2.5b is a zero-shot image classification model from cs-giung. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Contrastive Language-Image Pretraining (CLIP) model pre-trained on 2.5 billion data points of CommonCrawl at resolution 224x224. It was introduced in the paper Learning Transferable Visual Models From Natural Language…
Downloads · 30 days
4
2% of all-time downloads
All-time downloads
242
Public
Parameters
151M
605 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors605 MB · 100%
How the weights are stored.
F32151M · 100%
From the Hugging Face model README
Contrastive Language-Image Pretraining (CLIP) model pre-trained on 2.5 billion data points of CommonCrawl at resolution 224x224. It was introduced in the paper Learning Transferable Visual Models From Natural Language Supervision and further reproduced in the follow-up paper Demystifying CLIP Data.
The weights were converted from the b32_fullcc2.5b.pt file presented in the original repository.