Downloads · 30 days
67
3% of all-time downloads
moranyanuka/icc
icc is a text classification model from moranyanuka. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
The official checkpoint of ICC model, introduced in ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
Downloads · 30 days
67
3% of all-time downloads
All-time downloads
2.4K
Public
Parameters
82.1M
328 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors328 MB · 99%
From the Hugging Face model README
The official checkpoint of ICC model, introduced in ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
The ICC model is used to quantify the concreteness of image captions, and the intended use is finding the best captions in a noisy multimodal dataset. It can be achieved by simply running it over the captions and filtering out samples with low score. It works best in conjunction with CLIP based filtering.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("moranyanuka/icc")
model = AutoModelForSequenceClassification.from_pretrained("moranyanuka/icc").to("cuda")
captions = ["a great method of quantifying concreteness", "a man with a white shirt"]
text_ids = tokenizer(captions, padding=True, return_tensors="pt", truncation=True).to('cuda')
with torch.inference_mode():
icc_scores = model(**text_ids)['logits']
# tensor([[0.0339], [1.0068]])
</details>
bibtex:
@misc{yanuka2024icc,
title={ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation},
author={Moran Yanuka and Morris Alper and Hadar Averbuch-Elor and Raja Giryes},
year={2024},
eprint={2403.01306},
archivePrefix={arXiv},
primaryClass={cs.LG}
}