Downloads · 30 days
102
31% of all-time downloads
Futyn-Maker/ruscxn-embedder
ruscxn-embedder is a sentence similarity model from Futyn-Maker. Use it when you need a score for how close two texts are. It is set up for sentence-transformers.
This is a specialized sentence-transformers model fine-tuned from intfloat/multilingual-e5-large-instruct for finding Russian Constructicon patterns in text. The model is trained to compare Russian text examples with…
Downloads · 30 days
102
31% of all-time downloads
All-time downloads
326
Public
Parameters
560M
2.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
This is a specialized sentence-transformers model fine-tuned from intfloat/multilingual-e5-large-instruct for finding Russian Constructicon patterns in text. The model is trained to compare Russian text examples with construction patterns from the Russian Constructicon database, enabling semantic search for linguistic constructions.
This model is specifically designed to encode Russian text examples and Constructicon patterns into a shared embedding space where similar constructions are close together. It enables:
This model is designed to be used with the RusCxnPipe library for automatic Russian Constructicon pattern extraction:
from ruscxnpipe import SemanticSearch
# Initialize with this specific model
search = SemanticSearch(
model_name="Futyn-Maker/ruscxn-embedder",
query_prefix="Instruct: Given a sentence, find the constructions of the Russian Constructicon that it contains\nQuery: ",
pattern_prefix=""
)
# Find construction candidates
examples = ["Петр так и замер.", "Мы, мягко говоря, совсем не ладили."]
results = search.find_candidates(queries=examples, n=5)
for result in results:
print(f"Example: {result['query']}")
for candidate in result['candidates']:
print(f" Pattern: {candidate['pattern']} (similarity: {candidate['similarity']:.3f})")
For advanced users who want to use the model directly:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Futyn-Maker/ruscxn-embedder")
# Note: Use the correct prefixes for optimal performance
query_prefix = "Instruct: Given a sentence, find the constructions of the Russian Constructicon that it contains\nQuery: "
pattern_prefix = ""
# Encode a Russian example
example = query_prefix + "Петр так и замер."
example_embedding = model.encode(example)
# Encode construction patterns (no prefix needed)
patterns = [
"NP-Nom так и VP-Pfv",
"VP вокруг да около",
"мягко говоря, Cl"
]
pattern_embeddings = model.encode(patterns)
# Calculate similarities
from sentence_transformers.util import cos_sim
similarities = cos_sim(example_embedding, pattern_embeddings)
print(similarities)
While this model is optimized for Russian Constructicon pattern matching, it may also be useful for other tasks involving Russian linguistic patterns, such as:
However, performance on these tasks has not been systematically evaluated.
The model was trained on 15,298 examples from the Russian Constructicon database, where each training sample consists of:
The model was fine-tuned using CachedMultipleNegativesSymmetricRankingLoss to learn embeddings where:
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: XLMRobertaModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
The model achieved its best validation performance at epoch 5 with a validation loss of 0.1145.