Downloads · 30 days
0
xingslong/AdvFLYP
AdvFLYP is a zero-shot image classification model from xingslong. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product.
Official model checkpoints for Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models (CVPR Findings 2026).
Downloads · 30 days
0
Access
Public
Updated Jul 27, 2026
Repo size
1.4 GB
Likes
1
Public
Click a slice to open those files.
.tar1.4 GB · 100%
From the Hugging Face model README
Official model checkpoints for Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models (CVPR Findings 2026).
AdvFLYP adversarially finetunes CLIP ViT-B/32 on one million image-text pairs sampled from LAION-400M. The full variant adds logit- and feature-level regularisation.
| File | Variant | SHA-256 |
|---|---|---|
AdvFLYP_full_checkpoint.pth.tar | AdvFLYP with logit- and feature-level regularisation (AdvFLYP_full in the paper) | 97fe54264a98d09787a3e819405cf32e0d6e2c2eb0dd592746d42ec3974cac37 |
AdvFLYP_NonReg_checkpoint.pth.tar | AdvFLYP without regularisation | 263f07ba2612cdd96de9684b4ce822fddd545298e38ef7cbde21283cc1616377 |
The files are the original PyTorch training checkpoints released by the authors. They contain the finetuned partial state dict and training state expected by the official repository.
from huggingface_hub import hf_hub_download
checkpoint_path = hf_hub_download(
repo_id="xingslong/AdvFLYP",
filename="AdvFLYP_full_checkpoint.pth.tar",
)
print(checkpoint_path)
For the non-regularised checkpoint, set filename="AdvFLYP_NonReg_checkpoint.pth.tar".
Clone the official repository and prepare its environment and datasets as described there. Then run:
git clone https://github.com/Sxing2/AdvFLYP.git
cd AdvFLYP
TEST_MODEL_PATH=/path/to/AdvFLYP_full_checkpoint.pth.tar
TEST_SET=(cifar10 cifar100 STL10 Caltech101 Caltech256 oxfordpet flowers102 Food101 StanfordCars SUN397 Country211 fgvc_aircraft EuroSAT dtd)
python -m code.main \
--evaluate \
--root ./data \
--resume "$TEST_MODEL_PATH" \
--test_set "${TEST_SET[@]}" \
--batch_size 256 \
--test_attack_type pgd \
--test_eps 1 \
--test_numsteps 10 \
--test_stepsize 1
The official implementation initializes OpenAI CLIP ViT-B/32 and loads the checkpoint's partial_state_dict with strict=False.
The models were finetuned on one million reachable image-text pairs randomly sampled from LAION-400M. See the GitHub repository for the released data and preparation details.
These checkpoints are intended for research on zero-shot image classification and adversarial robustness. Evaluation in the paper covers 14 downstream datasets under the stated attack configuration. Performance outside those datasets, threat models, and preprocessing settings has not been established.
Because the training images originate from web-scale LAION data, the checkpoints may inherit biases and other limitations of CLIP and the underlying data.
@InProceedings{Xing_2026_CVPR,
author = {Xing, Songlong and Wang, Weijie and Zhao, Zhengyu and Gu, Jindong and Torr, Philip and Sebe, Nicu},
title = {Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings},
month = {June},
year = {2026},
pages = {737--747}
}