Downloads · 30 days
8
2% of all-time downloads
kwang2049/TSDAE-cqadupstack2nli_stsb
TSDAE-cqadupstack2nli_stsb is a feature extraction model from kwang2049. Use it when you need embeddings to search or compare text. It is set up for transformers.
This is a model from the paper "TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning". This model adapts the knowledge from the NLI and STSb data to the specifi…
Downloads · 30 days
8
2% of all-time downloads
All-time downloads
420
Public
Repo size
438 MB
Likes
0
Public
Click a slice to open those files.
.bin438 MB · 100%
From the Hugging Face model README
This is a model from the paper "TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning". This model adapts the knowledge from the NLI and STSb data to the specific domain cqadupstack. Training procedure of this model:
The pooling method is CLS-pooling.
To use this model, an convenient way is through SentenceTransformers. So please install it via:
pip install sentence-transformers
And then load the model and use it to encode sentences:
from sentence_transformers import SentenceTransformer, models
dataset = 'cqadupstack'
model_name_or_path = f'kwang2049/TSDAE-{dataset}2nli_stsb'
model = SentenceTransformer(model_name_or_path)
model[1] = models.Pooling(model[0].get_word_embedding_dimension(), pooling_mode='cls') # Note this model uses CLS-pooling
sentence_embeddings = model.encode(['This is the first sentence.', 'This is the second one.'])
To evaluate the model against the datasets used in the paper, please install our evaluation toolkit USEB:
pip install useb # Or git clone and pip install .
python -m useb.downloading all # Download both training and evaluation data
And then do the evaluation:
from sentence_transformers import SentenceTransformer, models
import torch
from useb import run_on
dataset = 'cqadupstack'
model_name_or_path = f'kwang2049/TSDAE-{dataset}2nli_stsb'
model = SentenceTransformer(model_name_or_path)
model[1] = models.Pooling(model[0].get_word_embedding_dimension(), pooling_mode='cls') # Note this model uses CLS-pooling
@torch.no_grad()
def semb_fn(sentences) -> torch.Tensor:
return torch.Tensor(model.encode(sentences, show_progress_bar=False))
result = run_on(
dataset,
semb_fn=semb_fn,
eval_type='test',
data_eval_path='data-eval'
)
Please refer to the page of TSDAE training in SentenceTransformers.
If you use the code for evaluation, feel free to cite our publication TSDAE: Using Transformer-based Sequential Denoising Auto-Encoderfor Unsupervised Sentence Embedding Learning:
@article{wang-2021-TSDAE,
title = "TSDAE: Using Transformer-based Sequential Denoising Auto-Encoderfor Unsupervised Sentence Embedding Learning",
author = "Wang, Kexin and Reimers, Nils and Gurevych, Iryna",
journal= "arXiv preprint arXiv:2104.06979",
month = "4",
year = "2021",
url = "https://arxiv.org/abs/2104.06979",
}