Downloads · 30 days
98
6% of all-time downloads
aiana94/NaSE
NaSE is a sentence similarity model from aiana94. Use it when you need a score for how close two texts are. It is set up for transformers. The card lists the license as apache-2.0.
This model is a news-adapted sentence encoder, domain-specialized starting from the pretrained massively mulitlingual sentence encoder LaBSE.
Downloads · 30 days
98
6% of all-time downloads
All-time downloads
1.7K
Public
Parameters
471M
5.7 GB on disk
Likes
3
Public
Click a slice to open those files.
.h51.9 GB · 33%
From the Hugging Face model README
This model is a news-adapted sentence encoder, domain-specialized starting from the pretrained massively mulitlingual sentence encoder LaBSE.
NaSE is a domain-adapted multilingual sentence encoder, initialized from LaBSE. It was specialized to the news domain using two multilingual corpora, namely Polynews and PolyNewsParallel. More specifically, NaSE was pretrained with two objectives: denoising auto-encoding and sequence-to-sequence machine translation.
Here is how to use this model to get the sentence embeddings of a given text in PyTorch:
from transformers import BertModel, BertTokenizerFast
tokenizer = BertTokenizerFast.from_pretrained('aiana94/NaSE')
model = BertModel.from_pretrained('aiana94/NaSE')
# pepare input
sentences = ["This is an example sentence", "Dies ist auch ein Beispielsatz in einer anderen Sprache."]
encoded_input = tokenizer(sentences, return_tensors='pt', padding=True)
# forward pass
with torch.no_grad():
output = model(**encoded_input)
# to get the sentence embeddings, use the pooler output
sentence_embeddings = output.pooler_output
and in Tensorflow:
from transformers import TFBertModel, BertTokenizerFast
tokenizer = BertTokenizerFast.from_pretrained('aiana94/NaSE')
model = TFBertModell.from_pretrained('aiana94/NaSE')
# pepare input
sentences = ["This is an example sentence", "Dies ist auch ein Beispielsatz in einer anderen Sprache."]
encoded_input = tokenizer(sentences, return_tensors='tf', padding=True)
# forward pass
with torch.no_grad():
output = model(**encoded_input)
# to get the sentence embeddings, use the pooler output
sentence_embeddings = output.pooler_output
For similarity between sentences, an L2-norm is recommended before calculating the similarity:
import torch
import torch.nn.functional as F
def cos_sim(a: torch.Tensor, b: torch.Tensor):
a_norm = F.normalize(a, p=2, dim=1)
b_norm = F.normalize(b, p=2, dim=1)
return torch.mm(a_norm, b_norm.transpose(0, 1))
Our model is intended to be used as a sentence, and in particular, news encoder. Given an input text, it outputs a vector which captures its semantic information. The sentence vector may be used for sentence similarity, information retrieval or clustering tasks.
NaSE was domain-adapted using two multilingual datasets: Polynews and the parallel PolyNewsParallel.
We use the following procedure to smoothen the per-language distribution when sampling for model training:
We initialize NaSE with the pretrained weights of the mulitlingual sentenece encoder LaBSE. Please refer to its model card or the corresponding paper for more detaled information about the pre-training procedure.
We adapt the multilingual sentence encoder to the news domain using two objectives:
NaSE is trained sequentially, first on reconstruction, and then on translation, i.e., we continue training the NaSE encoder obtained with the DAE objective for translation on parallel data.
The full training scripts is accessible in the training code.
The model was pretrained on 1 40GB NVIDIA A100 GPU for a total of 100k steps.
BibTeX:
@misc{iana2024news,
title={News Without Borders: Domain Adaptation of Multilingual Sentence Embeddings for Cross-lingual News Recommendation},
author={Andreea Iana and Fabian David Schmidt and Goran Glavaš and Heiko Paulheim},
year={2024},
eprint={2406.12634},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2406.12634}
}