Downloads · 30 days
0
chiyum609/ProtoViT
ProtoViT is a machine learning model from chiyum609. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains pretrained ProtoViT models for interpretable image classification, as described in our paper "Interpretable Image Classification with Adaptive Prototype-based Vision Transformers".
Downloads · 30 days
0
Access
Public
Updated Nov 2, 2024
Repo size
216 MB
Likes
2
Public
Click a slice to open those files.
.pth216 MB · 100%
From the Hugging Face model README
This repository contains pretrained ProtoViT models for interpretable image classification, as described in our paper "Interpretable Image Classification with Adaptive Prototype-based Vision Transformers".
ProtoViT combines Vision Transformers with prototype-based learning to create models that are both highly accurate and interpretable. Rather than functioning as a black box, ProtoViT learns interpretable prototypes that explain its classification decisions through visual similarities.
We provide three variants of ProtoViT:
All models were trained and evaluated on the CUB-200-2011 fine-grained bird species classification dataset.
| Model Version | Backbone | Resolution | Top-1 Accuracy | Checkpoint |
|---|---|---|---|---|
| ProtoViT-T | DeiT-Tiny | 224×224 | 83.36% | Download |
| ProtoViT-S | DeiT-Small | 224×224 | 85.30% | Download |
| ProtoViT-CaiT | CaiT_xxs24 | 224×224 | 86.02% | Download |
If you use this model in your research, please cite:
@article{ma2024interpretable,
title={Interpretable Image Classification with Adaptive Prototype-based Vision Transformers},
author={Ma, Chiyu and Donnelly, Jon and Liu, Wenjun and Vosoughi, Soroush and Rudin, Cynthia and Chen, Chaofan},
journal={arXiv preprint arXiv:2410.20722},
year={2024}
}
This implementation builds upon the following excellent repositories:
This project is released under [MIT] license.
For any questions or feedback, please: