Downloads · 30 days
65
3% of all-time downloads
nuvocare/WikiMedical_sent_bert
WikiMedical_sent_bert is a sentence similarity model from nuvocare. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as cc-by-4.0.
This is a sentence-transformers model: It maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search.
Downloads · 30 days
65
3% of all-time downloads
All-time downloads
2K
Public
Parameters
22.7M
182 MB on disk
Likes
0
Public
Click a slice to open those files.
.bin90.9 MB · 50%
From the Hugging Face model README
This is a sentence-transformers model: It maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search.
WikiMedical_sent_bert is based on the 'all-MiniLM-L6-v2' sentence-transformers backbone and has been trained on the WikiMedical_sentence_simialrity dataset.
The model is able to predict whether two texts are related to the same wikipedia page, with only medical topic.
Using this model becomes easy when you have sentence-transformers installed:
pip install -U sentence-transformers
Then you can use the model like this:
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer('WikiMedical_sent_bert')
embeddings = model.encode(sentences)
print(embeddings)
The model is evaluated on the test set of WikiMedical_sentence_similarity. It achieves a :
For an automated evaluation of this model, see the Sentence Embeddings Benchmark: https://seb.sbert.net
The model was trained with the parameters:
DataLoader:
torch.utils.data.dataloader.DataLoader of length 3170 with parameters:
{'batch_size': 16, 'sampler': 'torch.utils.data.sampler.RandomSampler', 'batch_sampler': 'torch.utils.data.sampler.BatchSampler'}
Loss:
sentence_transformers.losses.CosineSimilarityLoss.CosineSimilarityLoss
Parameters of the fit()-Method:
{
"epochs": 2,
"evaluation_steps": 0,
"evaluator": "sentence_transformers.evaluation.EmbeddingSimilarityEvaluator.EmbeddingSimilarityEvaluator",
"max_grad_norm": 1,
"optimizer_class": "<class 'torch.optim.adamw.AdamW'>",
"optimizer_params": {
"lr": 2e-05
},
"scheduler": "WarmupLinear",
"steps_per_epoch": null,
"warmup_steps": 300,
"weight_decay": 0.01
}
SentenceTransformer(
(0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False})
(2): Normalize()
)