Downloads · 30 days
15
41% of all-time downloads
zeromodels/beit-large-patch16-512
beit-large-patch16-512 is a image classification model from zeromodels. Use it when you need a label for an image. It is set up for zeromodels. The card lists the license as apache-2.0.
[](https://github.com/IMvision12/ZeroModels) [](https://imvision12.github.io/ZeroModels/beit/) [](https://huggingface.co/collections/zeromodels/beit-6a9352067192fd9fcfcfe6f1)
Downloads · 30 days
15
41% of all-time downloads
All-time downloads
37
Public
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.h51.2 GB · 100%
From the Hugging Face model README
Paper: BEiT: BERT Pre-Training of Image Transformers (arXiv:2106.08254) · HF Papers
BEiT is a ViT-family vision transformer with a per-layer relative position bias, a learnable layer scale on each residual branch, and mean pooling of the patch tokens. Large backbone fine-tuned on ImageNet-1k at 512x512 (1000 classes).
For more details on the model, please go to Microsoft's original model card.
Pure-Keras 3 conversion of microsoft/beit-large-patch16-512 for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.
This is a image classification checkpoint (BeitImageClassify).
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from PIL import Image
from zeromodels.models.beit import BeitImageClassify, BeitModel, BeitImageProcessor
model = BeitImageClassify.from_weights("zeromodels/beit-large-patch16-512")
processor = BeitImageProcessor.from_weights("zeromodels/beit-large-patch16-512")
image = Image.open("your_image.jpg").convert("RGB")
pixels = processor(image) # resize + normalize (normalization lives in the processor)
logits = model(pixels, training=False)
print(logits.shape) # (1, num_classes)
# Feature extraction: the backbone without the classifier head
backbone = BeitModel.from_weights("zeromodels/beit-large-patch16-512", as_backbone=True)
features = backbone(pixels, training=False)
Load any BEiT variant the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub | Task |
|---|---|---|
beit-base-patch16-224 | zeromodels/beit-base-patch16-224 | image classification |
beit-large-patch16-224 | zeromodels/beit-large-patch16-224 | image classification |
beit-large-patch16-512 | zeromodels/beit-large-patch16-512 | image classification |
beit-base-patch16-224-pt22k-ft22k | zeromodels/beit-base-patch16-224-pt22k-ft22k | image classification |
beit-large-patch16-224-pt22k-ft22k | zeromodels/beit-large-patch16-224-pt22k-ft22k | image classification |
beit-base-finetuned-ade-640-640 | zeromodels/beit-base-finetuned-ade-640-640 | semantic segmentation |
beit-large-finetuned-ade-640-640 | zeromodels/beit-large-finetuned-ade-640-640 | semantic segmentation |
KERAS_BACKEND before importing Keras / zeromodels.[0, 255] pixels.BeitImageClassify; semantic segmentation uses BeitSemanticSegment and returns logits at a quarter of the input resolution (upsample the argmax map to the input size).BeitModel.from_weights(..., as_backbone=True) returns the per-block token sequences for feature extraction.hf: prefix, e.g. BeitImageClassify.from_weights("hf:microsoft/beit-large-patch16-512").A huge thank you to the Microsoft Research BEiT authors for creating and releasing these models.
License: Apache 2.0.