Downloads · 30 days
27
11% of all-time downloads
radlab/semantic-euro-bert-encoder-v1
semantic-euro-bert-encoder-v1 is a sentence similarity model from radlab. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A Polish semantic embedder trained on pairs constructed from plWordNet (Słowosieć) semantic relations and external descriptions of meanings. Every relation between lexical units and synsets is transformed into trainin…
Downloads · 30 days
27
11% of all-time downloads
All-time downloads
254
Public
Parameters
608M
7.4 GB on disk
Likes
2
Public
Click a slice to open those files.
.pt4.9 GB · 66%
From the Hugging Face model README
A Polish semantic embedder trained on pairs constructed from plWordNet (Słowosieć) semantic relations and external descriptions of meanings. Every relation between lexical units and synsets is transformed into training/evaluation examples.
The dataset mixes meanings’ usage signals: emotions, definitions, and external descriptions (Wikipedia, sentence-split). The embedder mimics semantic relations: it pulls together embeddings that are linked by “positive” relations (e.g., synonymy, hypernymy/hyponymy as defined in the dataset) and pushes apart embeddings linked by “negative” relations (e.g., antonymy or mutually exclusive relations). Source code and training scripts:
sentence-transformers (transformer encoder + pooling).Constructed from plWordNet relations between lexical units and synsets; each relation yields example pairs. Augmented with:
Positive pairs correspond to relations expected to increase similarity; negative pairs correspond to relations expected to decrease similarity. Additional hard/soft negatives may include unrelated meanings.
SentenceTransformerTrainerCosineSimilarityLossEmbeddingSimilarityEvaluator (cosine)


Sentence-Transformers:
# Python
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("radlab/semantic-euro-bert-encoder-v1", trust_remote_code=True)
texts = ["zamek", "drzwi", "wiadro", "horyzont", "ocean"]
emb = model.encode(texts, convert_to_tensor=True, normalize_embeddings=True)
scores = util.cos_sim(emb, emb)
print(scores) # higher = more semantically similar
Transformers (feature extraction):
# Python
from transformers import AutoModel, AutoTokenizer
import torch
import torch.nn.functional as F
name = "radlab/semantic-euro-bert-encoder-v1"
tok = AutoTokenizer.from_pretrained(name)
mdl = AutoModel.from_pretrained(name, trust_remote_code=True)
texts = ["student", "żak"]
tokens = tok(texts, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
out = mdl(**tokens)
emb = out.last_hidden_state.mean(dim=1)
emb = F.normalize(emb, p=2, dim=1)
sim = emb @ emb.T
print(sim)