Downloads · 30 days
423
2% of all-time downloads
Seznam/simcse-small-e-czech
simcse-small-e-czech is a sentence similarity model from Seznam. Use it when you need a score for how close two texts are. It is set up for transformers. The card lists the license as cc-by-4.0.
SimCSE-Small-E-Czech is the Seznam/small-e-czech model fine-tuned with the SimCSE objective.
Downloads · 30 days
423
2% of all-time downloads
All-time downloads
24.4K
Public
Repo size
108 MB
Likes
1
Public
Click a slice to open those files.
.bin54 MB · 98%
From the Hugging Face model README
SimCSE-Small-E-Czech is the Seznam/small-e-czech model fine-tuned with the SimCSE objective.
This model was created at Seznam.cz as part of a project to create high-quality small Czech semantic embedding models. These models perform well across various natural language processing tasks, including similarity search, retrieval, clustering, and classification. For further details or evaluation results, please visit the associated paper or GitHub repository.
You can load and use the model like this:
import torch
from transformers import AutoModel, AutoTokenizer
model_name = "Seznam/retromae-small-cs" # Hugging Face link
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
input_texts = [
"Dnes je výborné počasí na procházku po parku.",
"Večer si oblíbím dobrý film a uvařím si čaj."
]
# Tokenize the input texts
batch_dict = tokenizer(input_texts, max_length=512, padding=True, truncation=True, return_tensors='pt')
outputs = model(**batch_dict)
embeddings = outputs.last_hidden_state[:, 0] # Extract CLS token embeddings
similarity = torch.nn.functional.cosine_similarity(embeddings[0], embeddings[1], dim=0)