Downloads · 30 days
151
67% of all-time downloads
gitmodelmujtaba/sapbert-snomed-loinc-rxnorm
sapbert-snomed-loinc-rxnorm is a feature extraction model from gitmodelmujtaba. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
[](https://huggingface.co/gitmodelmujtaba/sapbert-snomed-loinc-rxnorm) -brightgreen)
Downloads · 30 days
151
67% of all-time downloads
All-time downloads
226
Public
Parameters
109M
2.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.bin438 MB · 50%
From the Hugging Face model README
A specialized biomedical representation model, supervised Cross-Encoder Reranker, and Contrastive Metric Projection Adapter engineered for multi-ontology clinical entity resolution across:
Total Unified Vocabulary: Over 1,242,379 standardized clinical concepts.
0.1352 down to 0.1109 to aggressively separate confusing semantic neighbors.treats, caused_by, indicated_for, evaluates, anatomical_site), enabling sub-millisecond 1-hop lookups and multi-hop pathway discovery.import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
# 1. Load model and tokenizer
repo_id = "gitmodelmujtaba/sapbert-snomed-loinc-rxnorm"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModel.from_pretrained(repo_id)
model.eval()
# 2. Clinical mentions vs Standard Ontology Concepts
mentions = [
"heart attack",
"lap chole",
"sugar in urine",
]
concepts = [
"Acute myocardial infarction (disorder)",
"Laparoscopic cholecystectomy (procedure)",
"Glycosuria (finding)",
]
# 3. Generate 768-dim CLS embeddings
def get_embeddings(texts):
inputs = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
return F.normalize(outputs.last_hidden_state[:, 0, :], p=2, dim=1)
m_emb = get_embeddings(mentions)
c_emb = get_embeddings(concepts)
# 4. Cosine similarity matrix
similarity = torch.mm(m_emb, c_emb.T)
for i, mention in enumerate(mentions):
best_idx = similarity[i].argmax().item()
print(f"'{mention}' --> '{concepts[best_idx]}' (similarity: {similarity[i][best_idx]:.4f})")
The model's concepts interface directly with the 2.24M clinical relation graph for multi-hop clinical pathway discovery:
[1-Hop] laparoscopic cholecystectomy --[treats]--> gallstone pancreatitis (score: 0.88)
[2-Hop] laparoscopic cholecystectomy -[treats]-> abdominal pain -[caused_by]-> acute pancreatitis
[3-Hop] laparoscopic cholecystectomy -[treats]-> surgical -[caused_by]-> nausea -[caused_by]-> acute pancreatitis
gitmodelmujtaba)