Downloads · 30 days
310
100% of all-time downloads
Overwatch886/bge-base-agri
bge-base-agri is a sentence similarity model from Overwatch886. Use it when you need a score for how close two texts are. It is set up for sentence-transformers.
This is a sentence-transformers model finetuned from BAAI/bge-base-en-v1.5. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.
Downloads · 30 days
310
100% of all-time downloads
All-time downloads
310
Public
Parameters
109M
4.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt2.6 GB · 60%
From the Hugging Face model README
This is a sentence-transformers model finetuned from BAAI/bge-base-en-v1.5. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'cls', 'include_prompt': True})
(2): Normalize({})
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
'Represent this sentence for searching relevant passages: What fertiliser should I use for rice?',
'Crop Rice | Country Tanzania | Title Fertiliser recommendation for Rice | Text Fertiliser programme for rice (Oryza sativa). Apply phosphorus and potassium basally and split nitrogen into three: at transplanting/establishment, active tillering, and panicle initiation. Avoid blanket high nitrogen, which worsens blast and blight; aim for balanced ~100-120 kg N/ha by leaf-colour chart. | Source CGIAR',
'Crop Cowpea | Country Ghana | Title Fertiliser recommendation for Cowpea (Humid forest) | Text Fertiliser programme for cowpea in the humid forest zone, where heavy rainfall leaches and acidifies soils and fungal pressure is high. As a nitrogen-fixing legume, cowpea needs little nitrogen; apply a phosphorus-rich basal (SSP or DAP) to feed nodulation and add potassium on depleted soils. In the humid forest, correct soil acidity with lime, watch for rapid fungal spread under the canopy, and improve drainage on the heavy leached soils. | Source FAO',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8315, 0.4254],
# [0.8315, 1.0000, 0.5164],
# [0.4254, 0.5164, 1.0000]])
<!--
### Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details>
-->
<!--
### Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details>
-->
<!--
### Out-of-Scope Use
*List how the model may foreseeably be misused and address what users ought not to do with the model.*
-->
<!--
## Bias, Risks and Limitations
*What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
-->
<!--
### Recommendations
*What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
-->
| anchor | positive | negative | |
|---|---|---|---|
| type | string | string | string |
| details | <ul><li>min: 17 tokens</li><li>mean: 22.11 tokens</li><li>max: 29 tokens</li></ul> | <ul><li>min: 53 tokens</li><li>mean: 117.47 tokens</li><li>max: 275 tokens</li></ul> | <ul><li>min: 52 tokens</li><li>mean: 113.78 tokens</li><li>max: 269 tokens</li></ul> |
| anchor | positive | negative |
|---|---|---|
| <code>Represent this sentence for searching relevant passages: How do I cope with flooding and excess rain on my farm?</code> | <code>Crop (general) | Country Ghana | Title Adapting to flooding and excess rain (Highlands) | Text Adapting to flooding and excess rain in the cool highlands, where heavy dew and mild temperatures favour foliar fungal diseases. Plant on ridges or raised beds, open drainage channels, choose flood-tolerant varieties, and stagger planting to spread risk. In the highlands, the long dewy nights drive foliar disease, so prioritise resistant varieties, wider spacing for airflow, and early scouting. | Source AGRA</code> | <code>Crop (general) | Country Tanzania | Title Managing waterlogging (Sahel) | Text Managing waterlogging in the hot semi-arid Sahel, where rains are short and erratic and sandy soils leach nutrients quickly. Cut drainage channels, plant on ridges or raised beds, improve structure with organic matter, and choose tolerant crops/varieties. Because sandy Sahelian soils leach fast, split applications finely, micro-dose into the planting hole, and conserve every drop of moisture with mulch and tied ridges. | Source Plantwise</code> |
| <code>Represent this sentence for searching relevant passages: How do I cope with flooding and excess rain on my farm?</code> | <code>Crop (general) | Country Nigeria | Title Adapting to flooding and excess rain (Humid forest) | Text Adapting to flooding and excess rain in the humid forest zone, where heavy rainfall leaches and acidifies soils and fungal pressure is high. Plant on ridges or raised beds, open drainage channels, choose flood-tolerant varieties, and stagger planting to spread risk. In the humid forest, correct soil acidity with lime, watch for rapid fungal spread under the canopy, and improve drainage on the heavy leached soils. | Source FAO</code> | <code>Crop (general) | Country Nigeria | Title Managing waterlogging (Humid forest) | Text Managing waterlogging in the humid forest zone, where heavy rainfall leaches and acidifies soils and fungal pressure is high. Cut drainage channels, plant on ridges or raised beds, improve structure with organic matter, and choose tolerant crops/varieties. In the humid forest, correct soil acidity with lime, watch for rapid fungal spread under the canopy, and improve drainage on the heavy leached soils. | Source CGIAR</code> |
| <code>Represent this sentence for searching relevant passages: How do I cope with flooding and excess rain on my farm?</code> | <code>Crop (general) | Country Tanzania | Title Adapting to flooding and excess rain (Sahel) | Text Adapting to flooding and excess rain in the hot semi-arid Sahel, where rains are short and erratic and sandy soils leach nutrients quickly. Plant on ridges or raised beds, open drainage channels, choose flood-tolerant varieties, and stagger planting to spread risk. Because sandy Sahelian soils leach fast, split applications finely, micro-dose into the planting hole, and conserve every drop of moisture with mulch and tied ridges. | Source FAO</code> | <code>Crop (general) | Country Tanzania | Title Managing waterlogging (Guinea savanna) | Text Managing waterlogging in the Guinea savanna, with one long rainy season and moderately fertile soils. Cut drainage channels, plant on ridges or raised beds, improve structure with organic matter, and choose tolerant crops/varieties. In the Guinea savanna, time operations to the single rains, top-dress at the onset of steady rainfall, and use the long season for a full-duration variety. | Source ICRISAT</code> |
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false,
"directions": [
"query_to_doc"
],
"partition_mode": "joint",
"hardness_mode": null,
"hardness_strength": 0.0
}
per_device_train_batch_size: 32learning_rate: 2e-05weight_decay: 0.01warmup_steps: 0.1fp16: Truebatch_sampler: no_duplicatesdo_predict: Falseprediction_loss_only: Trueper_device_train_batch_size: 32per_device_eval_batch_size: 8gradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 2e-05weight_decay: 0.01adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_ratio: Nonewarmup_steps: 0.1log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Trueenable_jit_checkpoint: Falsesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseuse_cpu: Falseseed: 42data_seed: Nonebf16: Falsefp16: Truebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: -1ddp_backend: Nonedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonedisable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torch_fusedoptim_args: Nonegroup_by_length: Falselength_column_name: lengthproject: huggingfacetrackio_space_id: trackioddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Truepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_for_metrics: []eval_do_concat_batches: Trueauto_find_batch_size: Falsefull_determinism: Falseddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_num_input_tokens_seen: noneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Trueuse_cache: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}| Epoch | Step | Training Loss |
|---|---|---|
| 0.8333 | 10 | 1.3753 |
| 1.6667 | 20 | 0.6970 |
| 2.5 | 30 | 0.6573 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}
<!--
## Glossary
*Clearly define terms in order to be accessible across audiences.*
-->
<!--
## Model Card Authors
*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
-->
<!--
## Model Card Contact
*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
-->