Downloads · 30 days
94
3% of all-time downloads
microsoft/cvt-13-384
cvt-13-384 is a image classification model from microsoft. Use it when you need a label for an image. It is set up for transformers. The card lists the license as apache-2.0.
CvT-13 model pre-trained on ImageNet-1k at resolution 384x384. It was introduced in the paper CvT: Introducing Convolutions to Vision Transformers by Wu et al. and first released in this repository.
Downloads · 30 days
94
3% of all-time downloads
All-time downloads
3K
Public
Repo size
241 MB
Likes
0
Public
Click a slice to open those files.
.h580.7 MB · 50%
From the Hugging Face model README
CvT-13 model pre-trained on ImageNet-1k at resolution 384x384. It was introduced in the paper CvT: Introducing Convolutions to Vision Transformers by Wu et al. and first released in this repository.
Disclaimer: The team releasing CvT did not write a model card for this model so this model card has been written by the Hugging Face team.
Here is how to use this model to classify an image of the COCO 2017 dataset into one of the 1,000 ImageNet classes:
from transformers import AutoFeatureExtractor, CvtForImageClassification
from PIL import Image
import requests
url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
image = Image.open(requests.get(url, stream=True).raw)
feature_extractor = AutoFeatureExtractor.from_pretrained('microsoft/cvt-13-384')
model = CvtForImageClassification.from_pretrained('microsoft/cvt-13-384')
inputs = feature_extractor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
# model predicts one of the 1000 ImageNet classes
predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])