Downloads · 30 days
24
29% of all-time downloads
zeromodels/vit_base_patch16_384_augreg_in1k
vit_base_patch16_384_augreg_in1k is a image classification model from zeromodels. Use it when you need a label for an image. It is set up for zeromodels. The card lists the license as apache-2.0.
[](https://github.com/IMvision12/ZeroModels) [](https://imvision12.github.io/ZeroModels/classificationbackbones/) [](https://huggingface.co/collections/zeromodels/vit-6a8eae63ab67c7a19fb7a6da)
Downloads · 30 days
24
29% of all-time downloads
All-time downloads
83
Public
Repo size
348 MB
Likes
0
Public
Click a slice to open those files.
.h5348 MB · 100%
From the Hugging Face model README
Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv:2010.11929) · HF Papers
Vision Transformer (ViT) patches an image and runs a transformer encoder. Use ViTImageClassify for logits or ViTModel for tokens / per-block features via as_backbone=True.
For more details on the model, please go to the upstream model card.
Pure-Keras 3 conversion of timm/vit_base_patch16_384.augreg_in1k for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.
This is an image-classification / backbone checkpoint (ViTImageClassify / ViTModel).
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from PIL import Image
from zeromodels.models.vit import ViTImageClassify, ViTModel, ViTImageProcessor
model = ViTImageClassify.from_weights("zeromodels/vit_base_patch16_384_augreg_in1k")
processor = ViTImageProcessor.from_weights("zeromodels/vit_base_patch16_384_augreg_in1k")
image = Image.open("your_image.jpg").convert("RGB")
pixels = processor(image) # resize + normalize (normalization lives in the processor)
logits = model(pixels, training=False)
print(logits.shape) # (1, num_classes)
# Feature extraction: the backbone without the classifier head
backbone = ViTModel.from_weights("zeromodels/vit_base_patch16_384_augreg_in1k", as_backbone=True)
features = backbone(pixels, training=False)
Load any ViT variant the same way with from_weights("zeromodels/<variant>"):
KERAS_BACKEND before importing Keras / zeromodels.ViTImageClassify returns class logits; ViTModel returns features (as_backbone=True for multi-scale stages).ViTImageClassify.from_weights("hf:timm/vit_base_patch16_384.augreg_in1k").A huge thank you to the ViT authors and the timm / Hub communities for creating and releasing these models.
License: see YAML license (usually matches the upstream checkpoint).