Downloads · 30 days
0
billpsomas/convnext_small_dino_simpool_ep100
convnext_small_dino_simpool_ep100 is a image classification model from billpsomas. Use it when you need a label for an image. The card lists the license as cc-by-4.0.
ConvNeXt-S model with SimPool (gamma=2.0) trained on ImageNet-1k for 100 epochs. Self-supervision with DINO.
Downloads · 30 days
0
Access
Public
Updated Nov 30, 2023
Repo size
1.8 GB
Likes
0
Public
Click a slice to open those files.
.pth615 MB · 100%
From the Hugging Face model README
ConvNeXt-S model with SimPool (gamma=2.0) trained on ImageNet-1k for 100 epochs. Self-supervision with DINO.
SimPool is a simple attention-based pooling method at the end of network, introduced on this ICCV 2023 paper and released in this repository. Disclaimer: This model card is written by the author of SimPool, i.e. Bill Psomas.
Convolutional networks and vision transformers have different forms of pairwise interactions, pooling across layers and pooling at the end of the network. Does the latter really need to be different? As a by-product of pooling, vision transformers provide spatial attention for free, but this is most often of low quality unless self-supervised, which is not well studied. Is supervision really the problem?
SimPool is a simple attention-based pooling mechanism as a replacement of the default one for both convolutional and transformer encoders. For transformers, we completely discard the [CLS] token. Interestingly, we find that, whether supervised or self-supervised, SimPool improves performance on pre-training and downstream tasks and provides attention maps delineating object boundaries in all cases. One could thus call SimPool universal.
| k | top1 | top5 |
|---|---|---|
| 10 | 68.66 | 85.832 |
| 20 | 68.508 | 87.636 |
| 100 | 66.33 | 88.876 |
| 200 | 64.93 | 88.53 |
@misc{psomas2023simpool,
title={Keep It SimPool: Who Said Supervised Transformers Suffer from Attention Deficit?},
author={Bill Psomas and Ioannis Kakogeorgiou and Konstantinos Karantzalos and Yannis Avrithis},
year={2023},
eprint={2309.06891},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
@inproceedings{liu2022convnet,
title={A convnet for the 2020s},
author={Liu, Zhuang and Mao, Hanzi and Wu, Chao-Yuan and Feichtenhofer, Christoph and Darrell, Trevor and Xie, Saining},
booktitle={Proceedings of the IEEE/CVF conference on computer vision and pattern recognition},
pages={11976--11986},
year={2022}
}