Downloads · 30 days
16
100% of all-time downloads
fassabilf/sea-clip-tiny-abl-wit
sea-clip-tiny-abl-wit is a zero-shot image classification model from fassabilf. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for open_clip. The card lists the license as mit.
One row of the ablation table of SEA-CLIP-Tiny (ACCV 2026). Same architecture, pipeline and hyperparameters as the main model; the difference is the training mixture.
Downloads · 30 days
16
100% of all-time downloads
All-time downloads
16
Public
Repo size
185 MB
Likes
0
Public
Click a slice to open those files.
.bin185 MB · 100%
From the Hugging Face model README
One row of the ablation table of SEA-CLIP-Tiny (ACCV 2026). Same architecture, pipeline and hyperparameters as the main model; the difference is the training mixture.
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('hf-hub:fassabilf/sea-clip-tiny-abl-wit')
tokenizer = open_clip.get_tokenizer('hf-hub:fassabilf/sea-clip-tiny-abl-wit')
| Architecture | ViT-T/16 vision tower + 12-layer / 384-wide text tower, embed dim 512 |
| Tokenizer | CLIP BPE, vocab 49408, context length 77 |
| Parameters | 46.11M (5.62M vision + 40.49M text) |
| Training data | WIT (487K pairs) |
| Teacher | MetaCLIP2-ViT-B-16-worldwide |
Retrieval R@1 on the held-out splits of each training source, zero-shot ImageNet accuracy, and the paper's retrieval-only Avg@1 over XM3600, Flickr30k-200 and XTD-200 (%).
| CG R@1 | WIT R@1 | Bloom R@1 | ImageNet | R@1-Avg |
|---|---|---|---|---|
| 0.2 | 27.7 | 2.2 | 1.3 | 0.8 |
Training and evaluation code: https://github.com/fassabilf/sea-clip-tiny.
The exact training configuration of this checkpoint is in params.txt in this repo.
@inproceedings{seacliptiny2026,
title = {SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages},
booktitle = {Asian Conference on Computer Vision (ACCV)},
year = {2026}
}