Downloads · 30 days
15
17% of all-time downloads
keras/vit_base_patch16_224_imagenet
vit_base_patch16_224_imagenet is a machine learning model from keras. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for keras-hub.
Vision Transformer (ViT) adapts the Transformer architecture, originally designed for natural language processing, to the domain of computer vision. It treats images as sequences of patches, similar to how Transformer…
Downloads · 30 days
15
17% of all-time downloads
All-time downloads
88
Public
Repo size
2.1 GB
Likes
0
Public
Click a slice to open those files.
.h5690 MB · 100%
From the Hugging Face model README
Vision Transformer (ViT) adapts the Transformer architecture, originally designed for natural language processing, to the domain of computer vision. It treats images as sequences of patches, similar to how Transformers treat sentences as sequences of words.. It was introduced in the paper An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.
Keras and KerasHub can be installed with:
pip install -U -q keras-hub
pip install -U -q keras
| Model ID | img_size | Acc | Top-5 | Parameters |
|---|---|---|---|---|
| Base | ||||
| vit_base_patch16_224_imagenet | 224 | - | - | 85798656 |
| vit_base_patch_16_224_imagenet21k | 224 | - | - | 85798656 |
| vit_base_patch_16_384_imagenet | 384 | - | - | 86090496 |
| vit_base_patch32_224_imagenet21k | 224 | - | - | 87455232 |
| vit_base_patch32_384_imagenet | 384 | - | - | 87528192 |
| Large | ||||
| vit_large_patch16_224_imagenet | 224 | - | - | 303301632 |
| vit_large_patch16_224_imagenet21k | 224 | - | - | 303301632 |
| vit_large_patch16_384_imagenet | 224 | - | - | 303690752 |
| vit_large_patch32_224_imagenet21k | 224 | - | - | 305510400 |
| vit_large_patch32_384_imagenet | 224 | - | - | 305607680 |
| Huge | ||||
| vit_huge_patch14_224_imagenet21k | 224 | - | - | 630764800 |
image_classifier = keras_hub.models.ImageClassification.from_preset(
"vit_base_patch16_224_imagenet"
)
input_data = np.random.uniform(0, 1, size=(2, 224, 224, 3))
image_classifier(input_data)
backbone = keras_hub.models.Backbone.from_preset(
"vit_base_patch16_224_imagenet"
)
preprocessor = keras_hub.models.ViTImageClassifierPreprocessor.from_preset(
"vit_base_patch16_224_imagenet"
)
model = keras_hub.models.ViTImageClassifier(
backbone=backbone,
num_classes=len(CLASSES),
preprocessor=preprocessor,
)
image_classifier = keras_hub.models.ImageClassification.from_preset(
"hf://keras/vit_base_patch16_224_imagenet"
)
input_data = np.random.uniform(0, 1, size=(2, 224, 224, 3))
image_classifier(input_data)
backbone = keras_hub.models.Backbone.from_preset(
"hf://keras/vit_base_patch16_224_imagenet"
)
preprocessor = keras_hub.models.ViTImageClassifierPreprocessor.from_preset(
"hf://keras/vit_base_patch16_224_imagenet"
)
model = keras_hub.models.ViTImageClassifier(
backbone=backbone,
num_classes=len(CLASSES),
preprocessor=preprocessor,
)