Downloads · 30 days
75
26% of all-time downloads
cometadata/snowflake-arctic-ror-affiliations
snowflake-arctic-ror-affiliations is a sentence similarity model from cometadata. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A sentence embedding model fine-tuned for Research Organization Registry (ROR) affiliation matching.
Downloads · 30 days
75
26% of all-time downloads
All-time downloads
294
Public
Parameters
568M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
A sentence embedding model fine-tuned for Research Organization Registry (ROR) affiliation matching.
This model is fine-tuned from Snowflake/snowflake-arctic-embed-l-v2.0 using contrastive learning
on the AffilGood contrastive dataset. It produces embeddings optimized for matching affiliation
strings to ROR organization records.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("cometadata/snowflake-arctic-ror-affiliations")
# Encode affiliations
affiliations = [
"Department of Physics, MIT, Cambridge, MA",
"Harvard Medical School, Boston",
]
embeddings = model.encode(affiliations, normalize_embeddings=True)
# Encode ROR organization names for matching
organizations = [
"Massachusetts Institute of Technology",
"Harvard University",
]
org_embeddings = model.encode(organizations, normalize_embeddings=True)
# Compute similarity
import numpy as np
similarities = np.dot(embeddings, org_embeddings.T)
This model is designed for dense retrieval in affiliation matching pipelines. It should be used as the first-stage retriever to find candidate ROR organizations for a given affiliation string.
Fine-tuned on SIRIS-Lab/affilgood-contrastive-dataset, which contains 52,900 affiliation-organization pairs with curated hard negatives across 105 languages.
2026-01-07T08:08:33.561241+00:00