Downloads · 30 days
1.1K
5% of all-time downloads
DatologyAI/retr-opt-vit-b-32
retr-opt-vit-b-32 is a zero-shot image classification model from DatologyAI. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for open_clip. The card lists the license as apache-2.0.
DatologyAI CLIP Retrieval is a state-of-the-art contrastive vision-language model optimized for image-text retrieval tasks through advanced data curation. This retrieval-optimized ViT-B/32 model achieves competitive p…
Downloads · 30 days
1.1K
5% of all-time downloads
All-time downloads
21.5K
Public
Repo size
2.4 GB
Likes
9
Public
Click a slice to open those files.
.bin605 MB · 50%
From the Hugging Face model README
DatologyAI CLIP Retrieval is a state-of-the-art contrastive vision-language model optimized for image-text retrieval tasks through advanced data curation. This retrieval-optimized ViT-B/32 model achieves competitive performance with SigLIP2 while requiring significantly less compute.
DatologyAI's retrieval-optimized CLIP model demonstrates superior performance on retrieval benchmarks through targeted data curation strategies:
This model is optimized for image-text retrieval tasks, cross-modal search, and multimodal understanding applications.
import torch
from PIL import Image
import open_clip
# Load model and preprocessing
model, _, preprocess = open_clip.create_model_and_transforms('hf-hub:DatologyAI/retr-opt-vit-b-32')
tokenizer = open_clip.get_tokenizer('hf-hub:DatologyAI/retr-opt-vit-b-32')
# Load and process image
image = preprocess(Image.open("path/to/image.jpg")).unsqueeze(0)
# Define text candidates
texts = [
"a photo of a cat",
"a dog playing in the park",
"a beautiful sunset over the ocean",
"people walking in a city"
]
text_tokens = tokenizer(texts)
# Compute similarities
with torch.no_grad():
image_features = model.encode_image(image)
text_features = model.encode_text(text_tokens)
# Normalize features
image_features /= image_features.norm(dim=-1, keepdim=True)
text_features /= text_features.norm(dim=-1, keepdim=True)
# Calculate similarity
similarity = (100.0 * image_features @ text_features.T)
# Get top matches
values, indices = similarity[0].topk(len(texts))
for idx, score in zip(indices, values):
print(f"{texts[idx]}: {score.item():.2f}")
import torch
import open_clip
from typing import List
def retrieve_images(query: str, image_features: torch.Tensor, top_k: int = 5):
"""
Retrieve top-k images for a text query
Args:
query: Text description to search for
image_features: Pre-computed normalized image features [N, 512]
top_k: Number of images to retrieve
"""
# Encode text query
text_tokens = tokenizer([query])
with torch.no_grad():
text_features = model.encode_text(text_tokens)
text_features /= text_features.norm(dim=-1, keepdim=True)
# Compute similarities
similarities = (100.0 * text_features @ image_features.T).squeeze()
# Get top-k matches
values, indices = similarities.topk(top_k)
return indices.tolist(), values.tolist()
# Example usage
model, _, preprocess = open_clip.create_model_and_transforms('hf-hub:DatologyAI/retr-opt-vit-b-32')
tokenizer = open_clip.get_tokenizer('hf-hub:DatologyAI/retr-opt-vit-b-32')
# Pre-compute image features for your dataset
# image_features = ... # Shape: [num_images, 512]
# Search for images
indices, scores = retrieve_images("a red sports car", image_features)
DatologyAI's retrieval-optimized pipeline employs specialized curation techniques:
The model uses standard CLIP contrastive objectives without architectural modifications.
The model was trained on image-text pairs curated from the DataComp-XL dataset using DatologyAI's retrieval-optimized curation pipeline, selecting high-quality pairs that enhance cross-modal alignment.
| Benchmark | Metric | DatologyAI | SigLIP2 | MetaCLIP |
|---|---|---|---|---|
| MSCOCO | Retrieval@1 | 55.53% | 55.45% | 46.6% |
| Flickr30K | Retrieval@1 | 79.7% | 82.4% | 72.9% |
If you use this model, please cite:
@article{datologyai2025clip,
title={CLIP Gets a Data Upgrade: Outperforming SoTA with Improved Data Curation Only},
author={DatologyAI Team},
journal={DatologyAI Blog},
year={2025},
url={https://datologyai.com/blog/clip-data-upgrade}
}
For more details on our data curation methodology and comprehensive benchmark results, please visit our blog post.
Contact: [email protected]
DatologyAI Team - [email protected]