Downloads · 30 days
66
5% of all-time downloads
zer0int/CLIP-SAE-ViT-L-14
CLIP-SAE-ViT-L-14 is a zero-shot image classification model from zer0int. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
66
5% of all-time downloads
All-time downloads
1.4K
Public
Parameters
428M
6.4 GB on disk
Likes
32
Public
Click a slice to open those files.
.pt3.4 GB · 54%
From the Hugging Face model README
Love ❤️ this CLIP?
ᐅ Buy me a coffee on Ko-Fi ☕
<details> <summary>Or click here for address to send 🪙₿ BTC</summary>3PscBrWYvrutXedLmvpcnQbE12Py8qLqMK
</details>
SAE = Sparse autoencoder
Accuracy ImageNet/ObjectNet my GmP: 91% > SAE (this): 89% > OpenAI pre-trained: 84.5%
But, it's fun to use with e.g. Flux.1 - get the Text-Encoder TE only version ⬇️ and try it!
And this SAE CLIP has best results for linear probe @ LAION-AI/CLIP_benchmark (see below)
This CLIP direct download is also the best CLIP to use for HunyuanVideo.
Required: Use with my zer0int/ComfyUI-HunyuanVideo-Nyan node (changes influence of LLM vs. CLIP; otherwise, difference is very little).
<video controls autoplay src="https://cdn-uploads.huggingface.co/production/uploads/6490359a877fc29cb1b09451/g0vO1N4JalPp8oIAq5v38.mp4"></video>


