Downloads · 30 days
1.1K
16% of all-time downloads
antflydb/clipclap
clipclap is a feature extraction model from antflydb. Use it when you need embeddings to search or compare text. It is set up for onnxruntime. The card lists the license as mit.
CLIPCLAP is a unified multimodal embedding model that maps text, images, and audio into a shared 512-dimensional vector space. It combines OpenAI's CLIP (text + image) with LAION's CLAP (audio) through a trained linea…
Downloads · 30 days
1.1K
16% of all-time downloads
All-time downloads
7.1K
Public
Repo size
1 GB
Likes
2
Public
Click a slice to open those files.
.data887 MB · 86%
From the Hugging Face model README
CLIPCLAP is a unified multimodal embedding model that maps text, images, and audio into a shared 512-dimensional vector space. It combines OpenAI's CLIP (text + image) with LAION's CLAP (audio) through a trained linear projection.
Built by antflydb for use with Antfly Inference, a standalone ML inference service for embeddings, chunking, reranking, and local model serving.
Text ──→ CLIP text encoder ──→ text_projection ──→ 512-dim (CLIP space)
Image ──→ CLIP visual encoder ──→ visual_projection ──→ 512-dim (CLIP space)
Audio ──→ CLAP audio encoder ──→ audio_projection ──→ 512-dim (CLIP space)
openai/clip-vit-base-patch32).laion/larger_clap_music_and_speech. The audio projection combines CLAP's native audio projection (1024→512) with a trained 512→512 linear layer that maps CLAP audio space into CLIP space.All three modalities produce 512-dimensional L2-normalized embeddings that are directly comparable via cosine similarity.
# Pull and run the model
antfly inference pull antflydb/clipclap:gguf:Q4_K
antfly inference run
# Embed text
curl -X POST http://localhost:8082/embed \
-H "Content-Type: application/json" \
-d '{
"model": "clipclap",
"input": [
{"type": "text", "text": "a cat sitting on a windowsill"},
{"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}},
{"type": "audio_url", "audio_url": {"url": "https://example.com/cat-purring.wav"}}
]
}'
The audio projection layer bridges CLAP and CLIP embedding spaces. Training procedure:
The contrastive loss pushes matching audio-text pairs together while pushing non-matching pairs apart within each batch, preserving content discrimination.
| Parameter | Value |
|---|---|
| Training dataset | OpenSound/AudioCaps |
| Samples | 5000 audio-caption pairs |
| Epochs | 20 |
| Batch size | 256 |
| Learning rate | 1e-3 |
| Optimizer | Adam |
| Loss | Symmetric InfoNCE (temperature=0.07) |
| Train/val split | 90/10 |
| Component | Model |
|---|---|
| CLIP | openai/clip-vit-base-patch32 |
| CLAP | laion/larger_clap_music_and_speech |
| File | Description | Size |
|---|---|---|
text_model.onnx | CLIP text encoder | ~254 MB |
visual_model.onnx | CLIP visual encoder | ~330 MB |
text_projection.onnx | CLIP text projection (512→512) | ~4 KB |
visual_projection.onnx | CLIP visual projection (768→512) | ~6 KB |
audio_model.onnx | CLAP HTSAT audio encoder | ~590 MB |
audio_projection.onnx | Combined CLAP→CLIP projection (1024→512) | ~8 KB |
Additional files: clip_config.json, tokenizer.json, preprocessor_config.json, projection_training_metadata.json.
If you use CLIPCLAP, please cite the underlying models:
@inproceedings{radford2021clip,
title={Learning Transferable Visual Models From Natural Language Supervision},
author={Radford, Alec and Kim, Jong Wook and Hallacy, Chris and others},
booktitle={ICML},
year={2021}
}
@inproceedings{wu2023clap,
title={Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation},
author={Wu, Yusong and Chen, Ke and Zhang, Tianyu and others},
booktitle={ICASSP},
year={2023}
}