Downloads · 30 days
10
15% of all-time downloads
thearod5/pl-bert-siamese-encoder
pl-bert-siamese-encoder is a feature extraction model from thearod5. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
This repository contains the embedding model used to embed artifact for traceability link prediction.
Downloads · 30 days
10
15% of all-time downloads
All-time downloads
66
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
This repository contains the embedding model used to embed artifact for traceability link prediction.
used in the siamese models
This embedding model is the encoder portion of the siamese model used in the paper cited. This model utilized a relational classifier to create similarity scores between text pairs resembling a cross-encoder and consistently ranked almost as high as the top performer.
Used to embed software artifacts intended to be compared via cosine similarity.
Software traceability link prediction, Retrieval Augmented Generation, Artifact Clustering.
The intended vision for this model within a traceability link prediction pipeline, used to retrieve software artifacts for an LLM prompt, and for clustering.
This model could be used for a good set of starting weights for requirements classification.
This data uses open source git data which can be inaccurate and lead to unexpected results.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
parent_artifacts = [
"Display Artifacts",
]
texts = [
"Display Artifacts", // parent artifact
"A table view should be provided to display all project artifacts.", // child 1
"The system should be able to generate documentation for a set of artifacts." // child 2
]
embeddings = model.encode(texts, convert_to_tensor=False)
parent_embedding = embeddings[0:1]
children_embeddings = embeddings[1:]
# Compute cosine similarity
sim_matrix = cosine_similarity(parent_embedding, children_embeddings)
Please see cited paper for more information on training method, evaluation, and resuts.