Downloads · 30 days
270
66% of all-time downloads
kaustuk000/meridian
meridian is a machine learning model from kaustuk000. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
Hyperbolic multimodal retrieval built on frozen CLIP representations, combining Lorentz and Euclidean embedding spaces for compact and hierarchy-aware semantic search.
Downloads · 30 days
270
66% of all-time downloads
All-time downloads
410
Public
Parameters
41.5M
1.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors947 MB · 78%
From the Hugging Face model README
Hyperbolic multimodal retrieval built on frozen CLIP representations, combining Lorentz and Euclidean embedding spaces for compact and hierarchy-aware semantic search.
Meridian is a multimodal retrieval model built on top of CLIP ViT-B/16 that learns both hyperbolic (Lorentz) and Euclidean representations for images and text.
Unlike conventional retrieval systems that operate entirely in Euclidean space, Meridian maps semantic information onto a Lorentz manifold where hierarchical structure emerges naturally. This geometry enables cleaner semantic organization, improved hierarchy preservation, and compact multimodal representations while retaining strong retrieval performance.
Evaluation on MS-COCO retrieval.
| Variant | i2t R@1 | i2t R@5 | i2t R@10 | t2i R@1 | t2i R@5 | t2i R@10 |
|---|---|---|---|---|---|---|
| Meridian (64d) | 29.66 | 55.20 | 67.18 | 25.29 | 51.00 | 63.02 |
Where:
from transformers import AutoModel
model = AutoModel.from_pretrained(
"kaustuk000/meridian",
trust_remote_code=True,
)
model.eval()
The CLIP backbone is downloaded automatically.
The retrieval index loader is built directly into the model.
index = model.load_index("kaustuk000/meridian")
Available tensors:
index["tensors"]["h_image"]
index["tensors"]["e_image"]
index["tensors"]["h_text"]
index["tensors"]["e_text"]
Index tensors are stored on the Hub in FP16 format for storage efficiency and automatically converted to FP32 when loaded.
import torch
inputs = model.processor(
text=["a photo of a dog running on a beach"],
return_tensors="pt",
padding="max_length",
max_length=77,
)
eos_indices = inputs["attention_mask"].sum(dim=1) - 1
with torch.no_grad():
out = model.encode_text(
input_ids=inputs["input_ids"],
attention_mask=inputs["attention_mask"],
eos_indices=eos_indices,
)
h_text = out["h_text"]
e_text = out["e_text"]
These embeddings can be compared against:
index["tensors"]["h_text"]
index["tensors"]["e_text"]
from PIL import Image
import torch
image = Image.open("image.jpg").convert("RGB")
inputs = model.processor(
images=image,
return_tensors="pt",
)
with torch.no_grad():
out = model.encode_image(
pixel_values=inputs["pixel_values"]
)
h_image = out["h_image"]
e_image = out["e_image"]
These embeddings can be compared against:
index["tensors"]["h_image"]
index["tensors"]["e_image"]
Meridian consists of five major components:
OpenAI CLIP ViT-B/16 image encoder.
OpenAI CLIP text transformer.
Instead of relying solely on the final transformer layer, Meridian learns weighted combinations across all transformer blocks.
Two parallel embedding spaces are learned:
The hyperbolic branch maps features onto a Lorentz manifold using exponential-map operations.
A learned gating mechanism dynamically combines hyperbolic and Euclidean similarities for each sample.
The CLIP vision and text encoders remain completely frozen during training.
Meridian learns:
This enables improved hierarchical organization and retrieval behavior without modifying the pretrained CLIP backbone.
Meridian is designed to preserve hierarchical structure more effectively than conventional Euclidean embeddings.
The examples below compare hierarchical organization produced by the original frozen CLIP embeddings against the representations learned by Meridian. The CLIP vision and text encoders remain frozen throughout training; improvements arise from Meridian's learned layer aggregation, projection heads, and adaptive gating modules.
Mixed Animal Region
├── Dog
├── Cat
├── Flower
├── Lion
├── Tiger
├── Tree
└── Elephant
Feline Region
├── Domestic Cat
├── Tabby Cat
├── Kitten
├── Tiger
├── Tiger Cub
├── Lion
├── Lioness
└── Big Cat Cub
Meridian naturally organizes semantically related concepts into coherent neighborhoods while reducing cross-category mixing.
Hyperbolic space grows exponentially with distance from the origin.
This makes it particularly well suited for representing hierarchical data.
In Meridian:
The hyperbolic branch uses Lorentz manifold operations rather than standard Euclidean distance metrics.
The geodesic distance between two points is:
dL(x,y) = (1/√c) arcosh(-c⟨x,y⟩L)
where:
⟨x,y⟩L = -x₀y₀ + Σᵢ xᵢyᵢ
This geometry provides exponentially increasing representational capacity as embeddings move away from the origin.
The released checkpoint was trained using:
Meridian is suitable for:
GitHub:
https://github.com/kaustuk000/Meridian
@misc{singh2026meridian,
title = {Meridian: Hyperbolic Image--Text Representations},
author = {Kaustuk Pratap Singh},
year = {2026},
url = {https://github.com/kaustuk000/Meridian}
}
If you use concepts originating from MERU, please also cite:
@inproceedings{desai2023meru,
title = {Hyperbolic Image-Text Representations},
author = {Desai, Karan and Nickel, Maximilian and Rajpurohit, Tanmay and Johnson, Justin and Vedantam, Ramakrishna},
booktitle = {International Conference on Machine Learning},
year = {2023}
}
Meridian builds upon several foundational projects:
Special thanks to the authors of these projects for their contributions to multimodal representation learning.