Downloads · 30 days
10
1% of all-time downloads
ostris/CLIP-ViT-H-14-448
CLIP-ViT-H-14-448 is a zero-shot image classification model from ostris. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
You probably do not need this unless you are training your own IP Adapters.
Downloads · 30 days
10
1% of all-time downloads
All-time downloads
1.9K
Public
Parameters
633M
1.3 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
You probably do not need this unless you are training your own IP Adapters.
Modified version of the vision encoder of CLIP-ViT-H-14-laion2B-s32B-b79K to handle 448 x 448 inputs vs the original 224 x 224 inputs. It will probbaly not work for classification (as is), but will DIP work for for IP+ adapters that use CLIP-ViT-H, though they will need to be fine tuned a little more.
Hidden layer outputs go from (257, 1280) to (1025, 1280), which can be digested by the Resampler without modification or weight resizing.