Downloads · 30 days
314
39% of all-time downloads
BarraHome/vmware-embeddings-large-v1
vmware-embeddings-large-v1 is a sentence similarity model from BarraHome. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as mit.
A specialized sentence-transformers model fine-tuned for semantic search and information retrieval in technical documentation, with a focus on enterprise infrastructure and virtualization technologies.
Downloads · 30 days
314
39% of all-time downloads
All-time downloads
798
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
A specialized sentence-transformers model fine-tuned for semantic search and information retrieval in technical documentation, with a focus on enterprise infrastructure and virtualization technologies.
This model extends BAAI/bge-base-en-v1.5 with domain-specific fine-tuning for technical documentation retrieval. It generates 768-dimensional dense embeddings optimized for semantic similarity in enterprise technology contexts.
Primary Use Cases:
Optimized For:
This model is specialized for technical documentation and may not perform optimally for:
pip install sentence-transformers
from sentence_transformers import SentenceTransformer, util
# Load model
model = SentenceTransformer('BarraHome/vmware-embeddings-large-v1')
# Example queries and documents
queries = [
"How to configure high availability?",
"Steps to install guest tools"
]
documents = [
"High availability can be configured through the management interface...",
"To install guest tools, first mount the ISO image..."
]
# Generate embeddings
query_embeddings = model.encode(queries)
doc_embeddings = model.encode(documents)
# Calculate similarity
similarities = util.cos_sim(query_embeddings, doc_embeddings)
print(similarities)
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer('BarraHome/vmware-embeddings-large-v1')
# Your document corpus
corpus = [
"Documentation about high availability features...",
"Guide for load balancing configuration...",
"Instructions for live migration procedures..."
]
# Encode corpus
corpus_embeddings = model.encode(corpus, convert_to_tensor=True)
# Query
query = "How to enable high availability?"
query_embedding = model.encode(query, convert_to_tensor=True)
# Search
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=3)
# Display results
for hit in hits[0]:
print(f"Score: {hit['score']:.4f}")
print(f"Document: {corpus[hit['corpus_id']]}\n")
Evaluated on a held-out test set of 2,000 diverse technical queries:
| Metric | Base Model | Fine-tuned | Improvement |
|---|---|---|---|
| Recall@1 | 0.637 | 0.759 | +19.2% |
| Recall@3 | 0.805 | 0.927 | +15.2% |
| Recall@5 | 0.853 | 0.956 | +12.1% |
| Recall@10 | 0.906 | 0.979 | +8.0% |
| NDCG@10 | 0.775 | 0.879 | +13.4% |
The fine-tuned model shows consistent improvements across all metrics:
Detailed Metric Comparison:

Percentage Improvements:

SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': True})
(1): Pooling({'pooling_mode_cls_token': True})
(2): Normalize()
)
For optimal results:
Query Formulation:
Hybrid Search:
Batch Processing:
encode(..., batch_size=32) for large collectionsconvert_to_tensor=True for GPU accelerationReranking:
| Hardware | Batch Size | Throughput |
|---|---|---|
| RTX 3090 | 32 | ~850 docs/sec |
| A100 | 128 | ~2,100 docs/sec |
| CPU (16 cores) | 8 | ~180 docs/sec |
Minimum:
Recommended:
@misc{vmware-embeddings-2024,
author = {Alberto Ferrer},
title = {VMware Technical Documentation Embeddings},
year = {2024},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/BarraHome/vmware-embeddings-large-v1}}
}
@misc{bge-base-en-v1.5,
author = {BAAI},
title = {BGE Base English v1.5},
year = {2023},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/BAAI/bge-base-en-v1.5}}
}
MIT License
Copyright (c) 2024 [Your Name]
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Note: This model is intended for research and development. For production use, ensure compliance with your organization's policies and applicable regulations.