Downloads · 30 days
10
13% of all-time downloads
Slep/CondViT-B16-txt
CondViT-B16-txt is a feature extraction model from Slep. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Introduced in <a href=https://arxiv.org/abs/2306.02928LRVSF-Fashion: Extending Visual Search with Referring Instructions</a, Lepage et al. 2023
Downloads · 30 days
10
13% of all-time downloads
All-time downloads
80
Public
Parameters
1.3B
5.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.3 GB · 100%
From the Hugging Face model README
Introduced in <a href=https://arxiv.org/abs/2306.02928>**LRVSF-Fashion: Extending Visual Search with Referring Instructions*</a>, Lepage et al. 2023*
<div align="center"> <div id=links>| Data | Code | Models | Spaces |
|---|---|---|---|
| Full Dataset | Training Code | Categorical Model | LRVS-F Leaderboard |
| Test set | Benchmark Code | Textual Model | Demo |
Model finetuned from CLIP ViT-B/16 on LRVSF at 224x224. The conditioning text is preprocessed by a frozen Sentence T5-XL.
Research use only.
from PIL import Image
import requests
from transformers import AutoProcessor, AutoModel
import torch
model = AutoModel.from_pretrained("Slep/CondViT-B16-txt")
processor = AutoProcessor.from_pretrained("Slep/CondViT-B16-txt")
url = "https://huggingface.co/datasets/Slep/LAION-RVS-Fashion/resolve/main/assets/108856.0.jpg"
img = Image.open(requests.get(url, stream=True).raw)
txt = "a brown bag"
inputs = processor(images=[img], texts=[txt])
raw_embedding = model(**inputs)
normalized_embedding = torch.nn.functional.normalize(raw_embedding, dim=-1)