Downloads · 30 days
17
4% of all-time downloads
keras-io/cct
cct is a image classification model from keras-io. Use it when you need a label for an image. It is set up for tf-keras.
Based on the Compact Convolutional Transformers example on keras.io created by Sayak Paul.
Downloads · 30 days
17
4% of all-time downloads
All-time downloads
442
Public
Repo size
4.8 MB
Likes
0
Public
Click a slice to open those files.
.data-00000-of-000011.7 MB · 35%
From the Hugging Face model README
Based on the Compact Convolutional Transformers example on keras.io created by Sayak Paul.
As discussed in the Vision Transformers (ViT) paper, a Transformer-based architecture for vision typically requires a larger dataset than usual, as well as a longer pre-training schedule. ImageNet-1k (which has about a million images) is considered to fall under the medium-sized data regime with respect to ViTs. This is primarily because, unlike CNNs, ViTs (or a typical Transformer-based architecture) do not have well-informed inductive biases (such as convolutions for processing images). This begs the question: can't we combine the benefits of convolution and the benefits of Transformers in a single network architecture? These benefits include parameter-efficiency, and self-attention to process long-range and global dependencies (interactions between different regions in an image).
In Escaping the Big Data Paradigm with Compact Transformers, Hassani et al. present an approach for doing exactly this. They proposed the Compact Convolutional Transformer (CCT) architecture.
The model is trained using the CIFAR-10 dataset.
The following hyperparameters were used during training:
| name | learning_rate | decay | beta_1 | beta_2 | epsilon | amsgrad | weight_decay | exclude_from_weight_decay | training_precision |
|---|---|---|---|---|---|---|---|---|---|
| AdamW | 0.0010000000474974513 | 0.0 | 0.8999999761581421 | 0.9990000128746033 | 1e-07 | False | 9.999999747378752e-05 | None | float32 |
