Downloads ยท 30 days
81
31% of all-time downloads
YKYSpatz/ragproject_ver5
ragproject_ver5 is a sentence similarity model from YKYSpatz. Use it when you need a score for how close two texts are. It is set up for sentence-transformers.
This is a sentence-transformers model finetuned from YKYSpatz/ragprojectver2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paโฆ
Downloads ยท 30 days
81
31% of all-time downloads
All-time downloads
260
Public
Parameters
33.4M
133 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors133 MB ยท 99%
From the Hugging Face model README
This is a sentence-transformers model finetuned from YKYSpatz/ragproject_ver2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': True}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the ๐ค Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
"What were the findings of the patient's ocular examination?",
'The patient was diagnosed with a myopic fundus in the right eye with no breaks, chorioretinal atrophy, and a tilted disc with no visible cup. The left eye had a detached retina with PVR involving the macula.',
'The patient presented to the glaucoma services for assessment of anterior segment findings.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
<!--
### Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details>
-->
<!--
### Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details>
-->
<!--
### Out-of-Scope Use
*List how the model may foreseeably be misused and address what users ought not to do with the model.*
-->
<!--
## Bias, Risks and Limitations
*What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
-->
<!--
### Recommendations
*What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
-->
| sentence_0 | sentence_1 | sentence_2 | |
|---|---|---|---|
| type | string | string | string |
| details | <ul><li>min: 4 tokens</li><li>mean: 11.7 tokens</li><li>max: 33 tokens</li></ul> | <ul><li>min: 6 tokens</li><li>mean: 39.77 tokens</li><li>max: 161 tokens</li></ul> | <ul><li>min: 6 tokens</li><li>mean: 41.91 tokens</li><li>max: 199 tokens</li></ul> |
| sentence_0 | sentence_1 | sentence_2 |
|---|---|---|
| <code>How did the wound heal?</code> | <code>Wound checks were performed at one and six weeks following hospital discharge and the patient's wounds went on to heal without complications. No further intervention was required, and the patient made a good recovery.</code> | <code>Follow-up care was arranged with the primary care doctor as well as a surgical specialist to monitor the recovery and provide ongoing management. The patient was advised to follow a healthy diet and exercise regularly, and the surgical wound healing was monitored during follow-up appointments.</code> |
| <code>Laboratory test results for pycnodysostosis diagnosis</code> | <code>All laboratory investigations were normal, including liver and renal function tests, serum electrolytes, serum calcium, and phosphate. Alkaline phosphatase level was found to be within normal limits.</code> | <code>The study was complemented with magnetic resonance imaging of the sella turcica that confirmed the diagnosis of pituitary macroadenoma with apoplexy. Recommendation to follow up with a neurologist and an endocrinologist was also provided.</code> |
| <code>causes of dystocia in cattle</code> | <code>Dystocia due to uterine torsion resulting in stillbirth of a male calf with arthrogryposis, vertebral aplasia, and an abdominal midline defect.</code> | <code>The patient was diagnosed with pelvic inflammatory disease (PID) and perihepatitis resulting in cholecystitis caused by Chlamydia trachomatis, which constituted FHC syndrome.</code> |
{
"distance_metric": "TripletDistanceMetric.EUCLIDEAN",
"triplet_margin": 5
}
per_device_train_batch_size: 16per_device_eval_batch_size: 16num_train_epochs: 5multi_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 5max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin| Epoch | Step | Training Loss |
|---|---|---|
| 0.3687 | 500 | 4.7786 |
| 0.7375 | 1000 | 4.4425 |
| 1.1062 | 1500 | 4.4054 |
| 1.4749 | 2000 | 4.3814 |
| 1.8437 | 2500 | 4.3766 |
| 2.2124 | 3000 | 4.3568 |
| 2.5811 | 3500 | 4.3276 |
| 2.9499 | 4000 | 4.3506 |
| 3.3186 | 4500 | 4.3072 |
| 3.6873 | 5000 | 4.2919 |
| 4.0560 | 5500 | 4.268 |
| 4.4248 | 6000 | 4.2607 |
| 4.7935 | 6500 | 4.2515 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{hermans2017defense,
title={In Defense of the Triplet Loss for Person Re-Identification},
author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
year={2017},
eprint={1703.07737},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
<!--
## Glossary
*Clearly define terms in order to be accessible across audiences.*
-->
<!--
## Model Card Authors
*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
-->
<!--
## Model Card Contact
*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
-->