Downloads · 30 days
0
Potato-Scientist/multi-modal-search
multi-modal-search is a machine learning model from Potato-Scientist. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
A PyTorch-based multimodal model for image-text retrieval trained on COCO Captions dataset.
Downloads · 30 days
0
Access
Public
Updated Jan 29, 2026
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
.bin784 MB · 93%
From the Hugging Face model README
A PyTorch-based multimodal model for image-text retrieval trained on COCO Captions dataset.
Image Encoder: ViT-Base/16 (Vision Transformer)
Text Encoder: BERT-Base-Uncased
Projection: Linear layers project both modalities to 512-dim shared embedding space
Training Strategy: Frozen backbones with trainable projection heads
Evaluated on COCO val2017 (5,000 images, 25,000 captions):
| Metric | Score |
|---|---|
| Recall@5 | 21.95% |
| Recall@10 | 31.20% |
| Recall@20 | 42.50% |
import torch
from huggingface_hub import hf_hub_download
# Download model
model_path = hf_hub_download(
repo_id="Potato-Scientist/multi-modal-search",
filename="pytorch_model.bin"
)
# Load checkpoint
checkpoint = torch.load(model_path, map_location='cpu')
# Load into your model
from src.models.multimodal_model import MultimodalModel
model = MultimodalModel(
embedding_dim=512,
freeze_backbones=False,
pretrained=False
)
model.load_state_dict(checkpoint['model_state_dict'])
model.eval()
Try the live demo: Multimodal Search Demo
If you use this model, please cite:
@misc{multimodal-search-2026,
author = {Potato-Scientist},
title = {Multimodal Search Model},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/Potato-Scientist/multimodal-search-model}}
}