Downloads · 30 days
85
16% of all-time downloads
gowitheflow/LASER-cubed-bert-base-unsup
LASER-cubed-bert-base-unsup is a sentence similarity model from gowitheflow. Use it when you need a score for how close two texts are. It is set up for transformers.
Official model checkpoints of LA(SER)<sup3</sup (LASER-cubed) from EMNLP 2023 paper "Length is a Curse and a Blessing for Document-level Semantics"
Downloads · 30 days
85
16% of all-time downloads
All-time downloads
527
Public
Repo size
876 MB
Likes
2
Public
Click a slice to open those files.
.bin438 MB · 100%
From the Hugging Face model README
Official model checkpoints of LA(SER)<sup>3</sup> (LASER-cubed) from EMNLP 2023 paper "Length is a Curse and a Blessing for Document-level Semantics"
LASER-cubed-bert-base-unsup is an unsupervised model trained on wiki1M dataset. Without needing the training sets to have long texts, it provides surprising generalizability on long document retrieval.
Use the model with Sentence Transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("gowitheflow/LASER-cubed-bert-base-unsup")
text = "LASER-cubed is a dope model - It generalizes to long texts without needing the training sets to have long texts."
representation = model.encode(text)
Evaluate it with the BEIR framework:
from beir.retrieval import models
from beir.datasets.data_loader import GenericDataLoader
from beir.retrieval.evaluation import EvaluateRetrieval
from beir.retrieval.search.dense import DenseRetrievalExactSearch as DRES
# download the datasets with BEIR original repo youself first
data_path = './datasets/arguana'
corpus, queries, qrels = GenericDataLoader(data_folder=data_path).load(split="test")
model = DRES(models.SentenceBERT("gowitheflow/LASER-cubed-bert-base-unsup"), batch_size=512)
retriever = EvaluateRetrieval(model, score_function="cos_sim")
results = retriever.retrieve(corpus, queries)
ndcg, _map, recall, precision = retriever.evaluate(qrels, results, retriever.k_values)
Information Retrieval
The model is not for further fine-tuning to do other tasks (such as classification), as it's trained to do representation tasks with similarity matching.
max seq 256, batch size 128, lr 3e-05, 1 epoch, 10% warmup, 1 A100.
wiki 1M
Please refer to the paper.
BibTeX:
@inproceedings{xiao2023length,
title={Length is a Curse and a Blessing for Document-level Semantics},
author={Xiao, Chenghao and Li, Yizhi and Hudson, G and Lin, Chenghua and Al Moubayed, Noura},
booktitle={Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing},
pages={1385--1396},
year={2023}
}