Downloads · 30 days
120
20% of all-time downloads
EMBO/vicreg_our
vicreg_our is a feature extraction model from EMBO. Use it when you need embeddings to search or compare text. It is set up for sentence-transformers. The card lists the license as apache-2.0.
SODA-VEC embedding model trained with VICReg Our loss function. This model uses normalized embeddings with covariance, feature, and dot product losses (diagonal-only) to learn biomedical text representations.
Downloads · 30 days
120
20% of all-time downloads
All-time downloads
612
Public
Parameters
149M
13.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors596 MB · 99%
From the Hugging Face model README
SODA-VEC embedding model trained with VICReg Our loss function. This model uses normalized embeddings with covariance, feature, and dot product losses (diagonal-only) to learn biomedical text representations.
This model is part of the SODA-VEC (Scientific Open Domain Adaptation for Vector Embeddings) project, which focuses on creating high-quality embedding models for biomedical and life sciences text.
Key Features:
EMBO/soda-vec-data-full_pmc_title_abstract_pairedLoss Function: VICReg Our: normalized embeddings with covariance loss, feature loss, and dot product loss (diagonal-only)
We have implemented a series of changes from the original VICREG in the paper from Meta. Here we show the main differences:
| Feature | Original VICReg | VICReg Our | VICReg Our Contrast |
|---|---|---|---|
| Normalization | No | Yes (L2-normalized) | Yes (L2-normalized) |
| Invariance (MSE) | Yes | No | No |
| Variance (hinge) | Yes | No | No |
| Covariance | Yes (unnormalized) | Yes (normalized) | Yes (normalized) |
| Feature correlation | No | Yes (cross-view) | Yes (cross-view) |
| Sample similarity | No | Yes (diagonal only) | Yes (diagonal + off-diagonal) |
Coefficients: cov=1.0, feature=1.0, dot=1.0
Base Model: answerdotai/ModernBERT-base
Training Configuration:
Training Command:
python scripts/soda-vec-train.py --config vicreg_our --coeff_cov 1 --coeff_feature 1 --coeff_dot 1 --push_to_hub --hub_org EMBO --save_limit 5
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer("EMBO/vicreg_our")
# Encode sentences
sentences = [
"CRISPR-Cas9 gene editing in human cells",
"Genome editing using CRISPR technology"
]
embeddings = model.encode(sentences)
print(f"Embedding shape: {embeddings.shape}")
# Compute similarity
from sentence_transformers.util import cos_sim
similarity = cos_sim(embeddings[0], embeddings[1])
print(f"Similarity: {similarity.item():.4f}")
from transformers import AutoTokenizer, AutoModel
import torch
import torch.nn.functional as F
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("EMBO/vicreg_our")
model = AutoModel.from_pretrained("EMBO/vicreg_our")
# Encode sentences
sentences = [
"CRISPR-Cas9 gene editing in human cells",
"Genome editing using CRISPR technology"
]
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
# Mean pooling
embeddings = outputs.last_hidden_state.mean(dim=1)
# Normalize (for VICReg models)
embeddings = F.normalize(embeddings, p=2, dim=1)
# Compute similarity
similarity = F.cosine_similarity(embeddings[0:1], embeddings[1:2])
print(f"Similarity: {similarity.item():.4f}")
The model has been evaluated on comprehensive biomedical benchmarks including:
For detailed evaluation results, see the SODA-VEC benchmark notebooks.
This model is designed for:
If you use this model, please cite:
@software{soda_vec,
title = {SODA-VEC: Scientific Open Domain Adaptation for Vector Embeddings},
author = {EMBO},
year = {2024},
url = {https://github.com/source-data/soda-vec}
}
For questions or issues, please open an issue on the SODA-VEC GitHub repository.
Model Card Generated: 2025-11-10