Downloads ยท 30 days
10.7K
10% of all-time downloads
timm/PE-Core-B-16
PE-Core-B-16 is a zero-shot image classification model from timm. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for open_clip. The card lists the license as apache-2.0.
This is an OpenCLIP (image + text) remaped version of the the original
Downloads ยท 30 days
10.7K
10% of all-time downloads
All-time downloads
109K
Public
Repo size
3.6 GB
Likes
0
Public
Click a slice to open those files.
.bin1.8 GB ยท 50%
From the Hugging Face model README
This is an OpenCLIP (image + text) remaped version of the the original
[๐ Tech Report] [๐ PE Github (original weights)] [๐ OpenCLIP Github (these weights)]
Perception Encoder (PE) is a state-of-the-art encoder for image and video understanding trained via simple vision-language learning. It was introduced in "Perception Encoder: The best visual embeddings are not at the output of the network".
Model Developer: Meta
Model Overview: Perception Encoder (PE) is a family of large-scale vision encoder models with state-of-the-art performance on a large variety of vision tasks. By using a robust contrastive pretraining recipe and finetuning on synthetically aligned videos, PE not only outperforms all existing models on classification and retrieval, but it also internally produces strong, general features that scale for downstream tasks. PE unlocks the ability for large-scale contrastive pretraining to transfer to downstream tasks with alignment tuning to capitalize on those general features.
<img src="https://huggingface.co/facebook/PE-Core-G14-448/resolve/main/docs/pe_image1.png" style="width: 100%; margin: 0 auto; display: block;" />PE core is our base model trained with our robust image pretraining schedule and finetuned on the data generated by our synthetic video data engine.
PE core curently comes in 3 sizes. PE core G is the main checkpoint, with L and B models distilled from it.
| Scale | Tower | Params | Width | Depth | MLP | Heads | CLIP Dim | Resolution / Context Len |
|---|---|---|---|---|---|---|---|---|
| B/16 | Vision | 0.09B | 768 | 12 | 3072 | 12 | 1024 | 224px |
| Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | |
| L/14 | Vision | 0.32B | 1024 | 24 | 4096 | 16 | 1024 | 336px |
| Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | |
| G/14 | Vision | 1.88B | 1536 | 50 | 8960 | 16 | 1280 | 448px |
| Text | 0.47B | 1280 | 24 | 5120 | 20 | 1280 | 72 tokens |
All PE core models use an attention pooling block with 8 heads on top of the vision tower. The L and B models additionally have a class token for global aggregation. See the paper for more details.
PE core obtains extremely strong results across the board on zero-shot image classification and retrieval as well as zero-shot video classification and retrieval. We present a sample of its performance across those domains below.
| Model | Checkpoint | IN-1k | IN-v2 | IN-A | ObjectNet | COCO-T2I | Kinetics-400 | VTT-T2I |
|---|---|---|---|---|---|---|---|---|
| T/16 384px | PE-Core-T-16-384 | 62.1 | 54.7 | 21.1 | 43.9 | 33.0 | 41.5 | 28.8 |
| S/16 384px | PE-Core-S-16-384 | 72.7 | 65.0 | 49.5 | 60.0 | 42.6 | 55.0 | 39.3 |
| B/16 224px | PE-Core-B-16 | 78.4 | 71.7 | 62.4 | 71.9 | 50.9 | 65.6 | 47.6 |
| L/14 336px | PE-Core-L-14-336 | 83.5 | 77.9 | 89.0 | 84.7 | 57.1 | 73.4 | 50.3 |
| G/14 448px | PE-Core-bigG-14-448 | 85.4 | 80.2 | 92.6 | 88.2 | 58.1 | 76.9 | 51.2 |
PE core performs particularly well on the hard benchmarks such as ObjectNet and ImageNet-A.
import torch
from urllib.request import urlopen
from PIL import Image
import open_clip
model_id = 'hf-hub:timm/PE-Core-B-16'
model, _, preprocess = open_clip.create_model_and_transforms(model_id)
tokenizer = open_clip.get_tokenizer(model_id)
image = Image.open(urlopen(
'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png'
))
image = preprocess(image).unsqueeze(0)
labels_list = ["a dog", "a cat", "a donut", "a beignet"]
text = tokenizer(labels_list, context_length=model.context_length)
with torch.no_grad(), torch.cuda.amp.autocast():
image_features = model.encode_image(image, normalize=True)
text_features = model.encode_text(text, normalize=True)
text_probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)
zipped_list = list(zip(labels_list, [100 * round(p.item(), 3) for p in text_probs[0]]))
print("Label probabilities: ", zipped_list)
You can find more details in the original GitHub repo and the OpenCLIP repo.
If you find our code useful for your research, please consider citing:
@article{bolya2025PerceptionEncoder,
title={Perception Encoder: The best visual embeddings are not at the output of the network},
author={Daniel Bolya and Po-Yao Huang and Peize Sun and Jang Hyun Cho and Andrea Madotto and Chen Wei and Tengyu Ma and Jiale Zhi and Jathushan Rajasegaran and Hanoona Rasheed and Junke Wang and Marco Monteiro and Hu Xu and Shiyu Dong and Nikhila Ravi and Daniel Li and Piotr Doll{\'a}r and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}
@article{cho2025PerceptionLM,
title={PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding},
author={Jang Hyun Cho and Andrea Madotto and Effrosyni Mavroudi and Triantafyllos Afouras and Tushar Nagarajan and Muhammad Maaz and Yale Song and Tengyu Ma and Shuming Hu and Hanoona Rasheed and Peize Sun and Po-Yao Huang and Daniel Bolya and Suyog Jain and Miguel Martin and Huiyu Wang and Nikhila Ravi and Shashank Jain and Temmy Stark and Shane Moon and Babak Damavandi and Vivian Lee and Andrew Westbury and Salman Khan and Philipp Kr\"{a}henb\"{u}hl and Piotr Doll{\'a}r and Lorenzo Torresani and Kristen Grauman and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}