Downloads · 30 days
10
0% of all-time downloads
Derify/ChemRanker-alpha-qed-sim
ChemRanker-alpha-qed-sim is a text ranking model from Derify. Use it for the text ranking task on the model card, and read the license before you ship it in a product. It is set up for sentence-transformers. The card lists the license as apache-2.0.
This Cross Encoder is finetuned from Derify/ModChemBERT-IR-BASE using hard-negative triplets derived from Derify/pubchem10mgenmolsimilarity. Positive SMILES pairs are first filtered by quality and similarity constrain…
Downloads · 30 days
10
0% of all-time downloads
All-time downloads
2.8K
Public
Parameters
204M
408 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors408 MB · 100%
From the Hugging Face model README
This Cross Encoder is finetuned from Derify/ModChemBERT-IR-BASE using hard-negative triplets derived from Derify/pubchem_10m_genmol_similarity. Positive SMILES pairs are first filtered by quality and similarity constraints, then reduced to one strongest positive target per anchor molecule to create a high-signal training set for reranking. The model computes relevance scores for pairs of SMILES strings, enabling SMILES reranking and molecular semantic search.
For this variant, the positives are selected with a composite ranking criterion that combines high QED and similarity without an additional similarity-contribution cutoff. The quality stage uses strict inequality filtering (QED > 0.85, similarity > 0.5, with similarity also bounded below 1.0), and then keeps the top-scoring pair per anchor molecule.
Hard negatives are mined with Sentence Transformers using Derify/ChemMRL-beta as the teacher model and a TopK-PercPos-style margin setting based on NV-Retriever, with relative_margin=0.05 and max_negative_score_threshold = pos_score * percentage_margin. Training uses triplet-format samples with 5 mined negatives per anchor-positive pair and optimizes a multiple-negatives ranking objective, while reranking evaluation uses n-tuple samples with 30 mined negatives per query.
First install the Transformers and Sentence Transformers libraries:
pip install -U "transformers>=4.57.1,<5.0.0"
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import CrossEncoder
# Download from the 🤗 Hub
model = CrossEncoder("Derify/ChemRanker-alpha-qed-sim")
# Get scores for pairs of texts
pairs = [
['c1snnc1C[NH2+]Cc1cc2c(s1)CCC2', 'c1snnc1CCC[NH2+]Cc1cc2c(s1)CCC2'],
['c1sc2c(c1-c1nc(C3CCOC3)no1)CCCC2', 'O=C([O-])Cc1noc(-c2csc3c2CCCC3)n1'],
['c1sc(C[NH2+]C2CC2)nc1C[NH+]1CCN2CCCC2C1', 'c1sc(C[NH2+]C2CC2)nc1C1CC([NH+]2CCN3CCCC3C2)C1'],
['c1sc(CC[NH+]2CCOCC2)nc1C[NH2+]C1CC1', 'CCc1nc(C[NH2+]C2CC2)cs1'],
['c1sc(CC2CCC[NH2+]2)nc1C1CCCO1', 'c1sc(CC2CCC[NH2+]2)nc1C1CCCC1'],
]
scores = model.predict(pairs)
print(scores.shape)
# (5,)
# Or rank different texts based on similarity to a single text
ranks = model.rank(
'c1snnc1C[NH2+]Cc1cc2c(s1)CCC2',
[
'c1snnc1CCC[NH2+]Cc1cc2c(s1)CCC2',
'O=C([O-])Cc1noc(-c2csc3c2CCCC3)n1',
'c1sc(C[NH2+]C2CC2)nc1C1CC([NH+]2CCN3CCCC3C2)C1',
'CCc1nc(C[NH2+]C2CC2)cs1',
'c1sc(CC2CCC[NH2+]2)nc1C1CCCC1',
]
)
# [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]
<!--
### Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details>
-->
<!--
### Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details>
-->
<!--
### Out-of-Scope Use
*List how the model may foreseeably be misused and address what users ought not to do with the model.*
-->
{
"at_k": 10
}
| Metric | Value |
|---|---|
| map | 0.4266 |
| mrr@10 | 0.671 |
| ndcg@10 | 0.6901 |
| smiles_a | smiles_b | negative | |
|---|---|---|---|
| type | string | string | string |
| details | <ul><li>min: 19 characters</li><li>mean: 33.64 characters</li><li>max: 65 characters</li></ul> | <ul><li>min: 20 characters</li><li>mean: 34.24 characters</li><li>max: 48 characters</li></ul> | <ul><li>min: 19 characters</li><li>mean: 33.27 characters</li><li>max: 57 characters</li></ul> |
| smiles_a | smiles_b | negative |
|---|---|---|
| <code>c1sc2cc3c(cc2c1CC[NH2+]C1CC1)OCCO3</code> | <code>FC(F)(F)[NH2+]CCc1csc2cc3c(cc12)OCCO3</code> | <code>[NH3+]CCCc1cc2c(cc1C1CC1)OCO2</code> |
| <code>c1sc2cc3c(cc2c1CC[NH2+]C1CC1)OCCO3</code> | <code>FC(F)(F)[NH2+]CCc1csc2cc3c(cc12)OCCO3</code> | <code>COc1cc2c(cc1C[NH2+]C1CCC1)OCO2</code> |
| <code>c1sc2cc3c(cc2c1CC[NH2+]C1CC1)OCCO3</code> | <code>FC(F)(F)[NH2+]CCc1csc2cc3c(cc12)OCCO3</code> | <code>O=c1[nH]c2cc3c(cc2cc1CNC1CCCCC1)OCCO3</code> |
{
"scale": 10.0,
"num_negatives": 4,
"activation_fn": "torch.nn.modules.activation.Sigmoid"
}
| smiles_a | smiles_b | negative_1 | negative_2 | negative_3 | negative_4 | negative_5 | negative_6 | negative_7 | negative_8 | negative_9 | negative_10 | negative_11 | negative_12 | negative_13 | negative_14 | negative_15 | negative_16 | negative_17 | negative_18 | negative_19 | negative_20 | negative_21 | negative_22 | negative_23 | negative_24 | negative_25 | negative_26 | negative_27 | negative_28 | negative_29 | negative_30 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| type | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string | string |
| details | <ul><li>min: 17 characters</li><li>mean: 37.57 characters</li><li>max: 96 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 34.45 characters</li><li>max: 70 characters</li></ul> | <ul><li>min: 18 characters</li><li>mean: 35.67 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 12 characters</li><li>mean: 35.13 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.28 characters</li><li>max: 81 characters</li></ul> | <ul><li>min: 17 characters</li><li>mean: 35.36 characters</li><li>max: 73 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.12 characters</li><li>max: 70 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.09 characters</li><li>max: 84 characters</li></ul> | <ul><li>min: 16 characters</li><li>mean: 35.16 characters</li><li>max: 64 characters</li></ul> | <ul><li>min: 13 characters</li><li>mean: 35.26 characters</li><li>max: 90 characters</li></ul> | <ul><li>min: 11 characters</li><li>mean: 35.16 characters</li><li>max: 90 characters</li></ul> | <ul><li>min: 15 characters</li><li>mean: 35.36 characters</li><li>max: 74 characters</li></ul> | <ul><li>min: 12 characters</li><li>mean: 35.16 characters</li><li>max: 63 characters</li></ul> | <ul><li>min: 15 characters</li><li>mean: 35.51 characters</li><li>max: 73 characters</li></ul> | <ul><li>min: 13 characters</li><li>mean: 35.21 characters</li><li>max: 69 characters</li></ul> | <ul><li>min: 10 characters</li><li>mean: 34.93 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 17 characters</li><li>mean: 35.41 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 18 characters</li><li>mean: 35.1 characters</li><li>max: 72 characters</li></ul> | <ul><li>min: 15 characters</li><li>mean: 35.43 characters</li><li>max: 62 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.36 characters</li><li>max: 65 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.48 characters</li><li>max: 65 characters</li></ul> | <ul><li>min: 18 characters</li><li>mean: 35.25 characters</li><li>max: 65 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.48 characters</li><li>max: 81 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.38 characters</li><li>max: 68 characters</li></ul> | <ul><li>min: 18 characters</li><li>mean: 35.67 characters</li><li>max: 68 characters</li></ul> | <ul><li>min: 16 characters</li><li>mean: 35.53 characters</li><li>max: 67 characters</li></ul> | <ul><li>min: 14 characters</li><li>mean: 35.39 characters</li><li>max: 83 characters</li></ul> | <ul><li>min: 18 characters</li><li>mean: 35.74 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 17 characters</li><li>mean: 35.56 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 11 characters</li><li>mean: 35.37 characters</li><li>max: 64 characters</li></ul> | <ul><li>min: 16 characters</li><li>mean: 35.51 characters</li><li>max: 77 characters</li></ul> | <ul><li>min: 16 characters</li><li>mean: 35.35 characters</li><li>max: 72 characters</li></ul> |
| smiles_a | smiles_b | negative_1 | negative_2 | negative_3 | negative_4 | negative_5 | negative_6 | negative_7 | negative_8 | negative_9 | negative_10 | negative_11 | negative_12 | negative_13 | negative_14 | negative_15 | negative_16 | negative_17 | negative_18 | negative_19 | negative_20 | negative_21 | negative_22 | negative_23 | negative_24 | negative_25 | negative_26 | negative_27 | negative_28 | negative_29 | negative_30 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| <code>c1snnc1C[NH2+]Cc1cc2c(s1)CCC2</code> | <code>c1snnc1CCC[NH2+]Cc1cc2c(s1)CCC2</code> | <code>c1snnc1CCC[NH2+]Cc1cc2c(s1)CCC2</code> | <code>Cn1cc(C[NH2+]Cc2cc3c(s2)CCC3)nn1</code> | <code>Cn1cc(CC[NH2+]Cc2cc3c(s2)CCC3)nn1</code> | <code>Cc1cc(C[NH2+]Cc2csnn2)sc1C</code> | <code>NC(=O)c1csc(C[NH2+]Cc2cc3c(s2)CCC3)c1</code> | <code>Cc1cc(CC[NH2+]Cc2csnn2)sc1C</code> | <code>N#CCc1csc(C[NH2+]Cc2cc3c(s2)CCC3)c1</code> | <code>Ic1ccc(C[NH2+]Cc2cc3c(s2)CCC3)o1</code> | <code>c1ncc(C[NH2+]Cc2csnn2)s1</code> | <code>c1c(C[NH2+]CC2CC2)sc2c1CSCC2</code> | <code>N#Cc1cc(F)cc(C[NH2+]Cc2cc3c(s2)CCC3)c1</code> | <code>c1cc(C[NH2+]Cc2nc3c(s2)CCC3)no1</code> | <code>CCc1ccc(C[NH2+]Cc2csnn2)s1</code> | <code>NCc1csc(NCc2cc3c(s2)CCC3)n1</code> | <code>CNH+Cc1nnc(-c2cc3c(s2)CCCC3)o1</code> | <code>Fc1cc(C[NH2+]Cc2cc3c(s2)CCC3)ccc1Br</code> | <code>FC(F)(F)C[NH2+]Cc1cc2c(s1)CCSC2</code> | <code>c1cc(C[NH2+]Cc2cc3c(s2)CCC3)c[nH]1</code> | <code>Cc1cc(C)c(CC[NH2+]Cc2cc3c(s2)CCC3)c(C)c1</code> | <code>Oc1ccc(C[NH2+]Cc2cc3c(s2)CCC3)cc1Br</code> | <code>O=C([O-])c1ccc(CC[NH2+]Cc2cc3c(s2)CCC3)s1</code> | <code>c1c(C[NH2+]CC2CCCC2)sc2c1CCC2</code> | <code>O=C([O-])c1ccc(C[NH2+]Cc2cc3c(s2)CCC3)s1</code> | <code>COc1cc(C)cc(C[NH2+]Cc2cc3c(s2)CCC3)c1</code> | <code>CCc1cnc(C[NH2+]Cc2csnn2)s1</code> | <code>Clc1cc(C[NH2+]Cc2cc3c(s2)CCC3)ccc1Br</code> | <code>c1c(C[NH2+]CC2CC2)sc2c1CCCCC2</code> | <code>Cc1ccccc1C[NH2+]Cc1cc2c(s1)CCC2</code> | <code>c1cc(C[NH+]2CCCC2)sc1C[NH2+]Cc1cc2c(s1)CCC2</code> | <code>Cc1cc(C[NH2+]Cc2cc3c(s2)CCC3)ccc1F</code> |
| <code>c1sc2c(c1-c1nc(C3CCOC3)no1)CCCC2</code> | <code>O=C([O-])Cc1noc(-c2csc3c2CCCC3)n1</code> | <code>Nc1sc2c(c1-c1nc(C3CCOC3)no1)CCCC2</code> | <code>Nc1sc2c(c1-c1nc(C3CCC3)no1)CCCC2</code> | <code>c1c(-c2nc(C3CCCNC3)no2)sc2c1CCCCCC2</code> | <code>Nc1sccc1-c1nc(C2CCCOC2)no1</code> | <code>Nc1sc2c(c1-c1nc(C3CCCO3)no1)CCCC2</code> | <code>Cc1csc(-c2nc(C3CCOCC3)no2)c1N</code> | <code>Cc1oc2c(c1-c1nc(C3CCOC3)no1)C(=O)CCC2</code> | <code>c1c(-c2nc(C3C[NH2+]CCO3)no2)sc2c1CCCCC2</code> | <code>O=C([O-])Nc1sc2c(c1-c1nc(C3CC3)no1)CCCC2</code> | <code>c1cc2c(s1)CCCC2c1nc(C2CC2)no1</code> | <code>CC(=O)N1CCCC(c2noc(-c3cc4c(s3)CCCCCC4)n2)C1</code> | <code>Cc1cc(-c2nc([C@@H]3CCOC3)no2)c(N)s1</code> | <code>c1cc2c(nc1-c1noc(C3CCCOC3)n1)CCCC2</code> | <code>Nc1sccc1-c1nc(C2CCCC2)no1</code> | <code>c1cc2c(nc1-c1noc(C3CCOCC3)n1)CCCC2</code> | <code>[NH3+]C(c1noc(-c2cc3c(s2)CCCC3)n1)C1CC1</code> | <code>c1cc2c(c(-c3nc(C4CCOCC4)no3)c1)CCCN2</code> | <code>c1c(-c2nc(C3CC3)no2)nn2c1CCCC2</code> | <code>CN1CC(c2noc(-c3cc4c(s3)CCCC4)n2)CC1=O</code> | <code>O=C([O-])Cc1noc(-c2csc3c2CCCC3)n1</code> | <code>Oc1c(-c2nc(C3CCC(F)(F)C3)no2)ccc2c1CCCC2</code> | <code>Cc1cc(=O)c(-c2noc(C3CCCOC3)n2)c2n1CCC2</code> | <code>O=C([O-])CNc1sc2c(c1-c1nc(C3CC3)no1)CCCC2</code> | <code>CC1CCc2c(sc(N)c2-c2nc(C3CC3)no2)C1</code> | <code>Cn1nc(-c2nc(C3CCCO3)no2)c2c1CCCC2</code> | <code>O=C(Nc1sc2c(c1-c1nc(C3CC3)no1)COCC2)C1=CCCCC1</code> | <code>Cc1cscc1-c1noc(C2CCOCC2)n1</code> | <code>CC1(C)CCCc2sc(N)c(-c3nc(C4CC4)no3)c21</code> | <code>Clc1cc2c(c(-c3nc(C4CCOC4)no3)c1)OCC2</code> | <code>Nc1sc2c(c1-c1nnc(C3CC3)o1)CCCC2</code> |
| <code>c1sc(C[NH2+]C2CC2)nc1C[NH+]1CCN2CCCC2C1</code> | <code>c1sc(C[NH2+]C2CC2)nc1C1CC([NH+]2CCN3CCCC3C2)C1</code> | <code>c1sc(C[NH2+]C2CC2)nc1C1CC([NH+]2CCN3CCCC3C2)C1</code> | <code>CC(C)[NH2+]Cc1nc(C[NH+]2CCC3CCCCC3C2)cs1</code> | <code>CN1C2CCC1CNH+CC2</code> | <code>Nc1nc(CC[NH+]2CCCN3CCCC3C2)cs1</code> | <code>CC1CNH+CCN1C</code> | <code>Oc1csc(CN2CCCC3C[NH2+]CC32)n1</code> | <code>CCc1nc(C[NH+]2CCCC3CCCCC32)cs1</code> | <code>C[NH2+]Cc1csc(N2CC[NH+]3CCCC3C2)n1</code> | <code>[NH3+]Cc1nc(C[NH+]2CCC3CCCCC32)cs1</code> | <code>CC1CN2CCCCC2C[NH+]1Cc1csc(CC[NH3+])n1</code> | <code>CCCc1nc(CN2CCCC2C2CCC[NH2+]2)cs1</code> | <code>ClCCc1nc(CN2CCCC2C2CCC[NH2+]2)cs1</code> | <code>c1cc(C[NH2+]C2CC2)c(C[NH+]2CCN3CCCCC3C2)o1</code> | <code>O=C(Cc1nc(CCl)cs1)N1CCC[NH+]2CCCC2C1</code> | <code>N#CCc1nc(C[NH+]2CCCC3CCCCC32)cs1</code> | <code>CC[NH2+]Cc1csc(N2CCC3C(CCC[NH+]3C)C2)n1</code> | <code>c1sc(C[NH2+]C2CC2)nc1C[NH+]1CCCCC1</code> | <code>[NH3+]Cc1nc(C[NH+]2CCCC2C2CCCC2)cs1</code> | <code>Cc1csc(C[NH+]2CCC3C[NH2+]CC3C2)n1</code> | <code>ClOCc1csc(C[NH+]2CC3C[NH2+]CC3C2)n1</code> | <code>c1cc(C[NH+]2CCCN3CCCC3C2)nc(C2CC2)n1</code> | <code>Cc1ccsc1C[NH2+]CCN1CCN2CCCC2C1</code> | <code>c1sc(C[NH2+]C2CCCC2)nc1C[NH+]1CCCCC1</code> | <code>Brc1csc(C[NH2+]CCN2CCN3CCCCC3C2)c1</code> | <code>Cc1nc(CCC[NH2+]C2CCN3CCCCC23)cs1</code> | <code>CCOC(=O)c1nc(CN2CC3CCC[NH2+]C3C2)cs1</code> | <code>CCCC(=O)c1nc(CN2CC3CCC[NH2+]C3C2)cs1</code> | <code>CC(C)(C)c1csc(CN2CCC[NH2+]C(C3CC3)C2)n1</code> | <code>COCc1nc(CN2CCC([NH3+])C2)cs1</code> | <code>CCC[NH2+]Cc1nc(C[NH+]2CC3CCC2C3)cs1</code> |
{
"scale": 10.0,
"num_negatives": 4,
"activation_fn": "torch.nn.modules.activation.Sigmoid"
}
eval_strategy: epochper_device_train_batch_size: 256per_device_eval_batch_size: 256torch_empty_cache_steps: 1000learning_rate: 3e-05weight_decay: 1e-05max_grad_norm: Nonelr_scheduler_type: warmup_stable_decaylr_scheduler_kwargs: {'num_decay_steps': 6238, 'warmup_type': 'linear', 'decay_type': '1-sqrt'}warmup_steps: 6238seed: 12data_seed: 24681357bf16: Truebf16_full_eval: Truetf32: Truedataloader_num_workers: 8dataloader_prefetch_factor: 2load_best_model_at_end: Trueoptim: stable_adamwoptim_args: decouple_lr=True,max_lr=3e-05dataloader_persistent_workers: Trueresume_from_checkpoint: Falsegradient_checkpointing: Truetorch_compile: Truetorch_compile_backend: inductortorch_compile_mode: max-autotuneeval_on_start: Truebatch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: epochprediction_loss_only: Trueper_device_train_batch_size: 256per_device_eval_batch_size: 256per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: 1000learning_rate: 3e-05weight_decay: 1e-05adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: Nonenum_train_epochs: 3max_steps: -1lr_scheduler_type: warmup_stable_decaylr_scheduler_kwargs: {'num_decay_steps': 6238, 'warmup_type': 'linear', 'decay_type': '1-sqrt'}warmup_ratio: 0.0warmup_steps: 6238log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 12data_seed: 24681357jit_mode_eval: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Truefp16_full_eval: Falsetf32: Truelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Truedataloader_num_workers: 8dataloader_prefetch_factor: 2past_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: stable_adamwoptim_args: decouple_lr=True,max_lr=3e-05adafactor: Falsegroup_by_length: Falselength_column_name: lengthproject: huggingfacetrackio_space_id: trackioddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Trueskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Falsehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Truegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Truetorch_compile_backend: inductortorch_compile_mode: max-autotuneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: noneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Trueuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Trueprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}| Epoch | Step | Training Loss | Validation Loss | ndcg@10 |
|---|---|---|---|---|
| 0.0002 | 1 | 1.2724 | - | - |
| 0.1603 | 1000 | 0.1583 | - | - |
| 0.3206 | 2000 | 0.0196 | - | - |
| 0.4809 | 3000 | 0.0112 | - | - |
| 0.6412 | 4000 | 0.0079 | - | - |
| 0.8015 | 5000 | 0.0063 | - | - |
| 0.9618 | 6000 | 0.0053 | - | - |
| 1.0 | 6238 | - | 1.6835 | 0.6811 |
| 1.1222 | 7000 | 0.0045 | - | - |
| 1.2825 | 8000 | 0.0041 | - | - |
| 1.4428 | 9000 | 0.0037 | - | - |
| 1.6031 | 10000 | 0.0034 | - | - |
| 1.7634 | 11000 | 0.0032 | - | - |
| 1.9237 | 12000 | 0.003 | - | - |
| 2.0 | 12476 | - | 1.6853 | 0.6891 |
| 2.0840 | 13000 | 0.0028 | - | - |
| 2.2443 | 14000 | 0.0026 | - | - |
| 2.4046 | 15000 | 0.0026 | - | - |
| 2.5649 | 16000 | 0.0025 | - | - |
| 2.7252 | 17000 | 0.0024 | - | - |
| 2.8855 | 18000 | 0.0023 | - | - |
| 3.0 | 18714 | - | 1.6982 | 0.6901 |
Carbon emissions were measured using CodeCarbon.
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{moreira2025nvretrieverimprovingtextembedding,
title={NV-Retriever: Improving text embedding models with effective hard-negative mining},
author={Gabriel de Souza P. Moreira and Radek Osmulski and Mengyao Xu and Ronay Ak and Benedikt Schifferer and Even Oldridge},
year={2025},
eprint={2407.15831},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2407.15831},
}
<!--
## Glossary
*Clearly define terms in order to be accessible across audiences.*
-->
<!--
## Model Card Authors
*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
-->
<!--
## Model Card Contact
*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
-->