Downloads · 30 days
11
3% of all-time downloads
asnelt/mnistvit
mnistvit is a machine learning model from asnelt. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
A vision transformer (ViT) trained on MNIST with a PyTorch-only implementation, achieving 99.65% test set accuracy.
Downloads · 30 days
11
3% of all-time downloads
All-time downloads
323
Public
Repo size
111 MB
Likes
1
Public
Click a slice to open those files.
.pt22.3 MB · 100%
From the Hugging Face model README
A vision transformer (ViT) trained on MNIST with a PyTorch-only implementation, achieving 99.65% test set accuracy.
The model is a vision transformer, as described in the original Dosovitskiy et al., ICLR 2021 paper.
The model is intended to be used for learning about vision transformers. It is small and trained on MNIST as a simple and well understood dataset. Together with the mnistvit package code, the importance of various hyperparameters can be explored.
Install the mnistvit package, which provides code for training and running the model:
pip install mnistvit
Place the config.json and model.pt file from this repository in a directory of your
choice and run Python from that directory.
To evaluate the test set accuracy and loss of the model stored in model.pt with
configuration config.json:
python -m mnistvit --use-accuracy --use-loss
Individual images can be classified as well. To predict the class of a digit image
stored in a file sample.jpg:
python -m mnistvit --image-file sample.jpg
This model was trained on the 60,000 training set images of the
MNIST dataset. Data augmentation was
used in the form of random rotations, translations and scaling as detailed in the
mnistvit.preprocess module.
Hyperparameters were obtained from an 80:20 training set - validation set split of the
original MNIST training set, running Ray Tune with Optuna as detailed in the
mnistvit.tune module. The resulting parameters were then set as default parameters in
the mnistvit.train module.
This model was evaluated on the 10,000 test set images of the MNIST dataset.
Test set accuracy: 99.65%
Test set cross entropy loss: 0.011