Downloads · 30 days
8
3% of all-time downloads
keras-io/convmixer
convmixer is a machine learning model from keras-io. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tf-keras. The card lists the license as apache-2.0.
The ConvMixer model is trained on Cifar10 dataset and is based on the paper, github.
Downloads · 30 days
8
3% of all-time downloads
All-time downloads
285
Public
Repo size
1.4 MB
Likes
0
Public
Click a slice to open those files.
.data-00000-of-000017.2 MB · 83%
From the Hugging Face model README
The ConvMixer model is trained on Cifar10 dataset and is based on the paper, github.
Disclaimer : This is a demo model for Sayak Paul's keras example. Please refrain from using this model for any other purpose.
The paper uses 'patches' (square group of pixels) extracted from the image, which has been done in other Vision Transformers like ViT. One notable dawback of such architectures is the quadratic runtime of self-attention layers which takes a lot of time and resources to train for usable output. The ConvMixer model, instead uses Convolutions along with the MLP-mixer to obtain similar results to that of transformers at a fraction of cost.
This model is intended to be used as a demo model for keras-io.