Downloads · 30 days
22
47% of all-time downloads
victormuryn/use-generated-pt
use-generated-pt is a sentence similarity model from victormuryn. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of paraphrase-multilingual-mpnet-base-v2, trained on the Ukrainian text corpus UberText 2.0 with generated augmentation with pool targets. It is part of the Ukrainian Sentence Embedd…
Downloads · 30 days
22
47% of all-time downloads
All-time downloads
47
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
This model is a fine-tuned version of paraphrase-multilingual-mpnet-base-v2, trained on the Ukrainian text corpus UberText 2.0 with generated augmentation with pool targets. It is part of the Ukrainian Sentence Embeddings collection, which explores the effect of different training strategies on sentence embedding quality for Ukrainian.
The model was fine-tuned using a contrastive objective on UberText 2.0, generated augmentation.
| Model | Description |
|---|---|
| use-natural-no-pt | Raw UberText 2.0, no augmentation, no pool targets |
| use-natural-pt | Raw UberText 2.0, no augmentation, pool targets |
| use-generated-no-pt | Generated augmentation, no pool targets |
| use-generated-pt | Generated augmentation, pool targets |
| use-translation-no-pt | Back-translation augmentation, no pool targets |
| use-translation-pt | Back-translation augmentation, pool targets |
| use-mask-no-pt | Masking augmentation, no pool targets |
| use-mask-pt | Masking augmentation, pool targets |
| use-dropout-no-pt | Dropout augmentation, no pool targets |
| use-dropout-pt | Dropout augmentation, pool targets |
| use-token-shuffle-no-pt | Token-shuffling augmentation, no pool targets |
| use-token-shuffle-pt | Token-shuffling augmentation, pool targets |
| use-combined-no-pt | Combined augmentation strategies, no pool targets |
| use-combined-pt | Combined augmentation strategies, pool targets |
| use-stochastic-no-pt | Markov-based stochastic augmentation, no pool targets |
| use-stochastic-pt | Markov-based stochastic augmentation, pool targets |
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("victormuryn/use-generated-pt")
sentences = [
"Проводжає сина мати захищати рідний край",
"Хоч би малесеньку хатину він мріяв мати над Дніпром",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
To be added
Apache 2.0