Downloads · 30 days
12
3% of all-time downloads
rbroc/contrastive-user-encoder-multipost
contrastive-user-encoder-multipost is a feature extraction model from rbroc. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
This model is a DistilBertModel trained by fine-tuning distilbert-base-uncased on author-based triplet loss.
Downloads · 30 days
12
3% of all-time downloads
All-time downloads
427
Public
Repo size
531 MB
Likes
0
Public
Click a slice to open those files.
.bin265 MB · 100%
From the Hugging Face model README
This model is a DistilBertModel trained by fine-tuning distilbert-base-uncased on author-based triplet loss.
Training and evaluation details are provided in our EMNLP Findings paper:
We fine-tuned DistilBERT on triplets consisting of:
rbroc/contrastive-user-encoder-singlepost for an equivalent model trained on a single anchor;To compute the loss, we use [CLS] encodings of the anchors, positive examples and negative examples from the last layer of the DistilBERT encoder. We perform feature-wise averaging of anchor posts encodings and optimize for \(max(||\overline{f(A)} - f(n)|| - ||\overline{f(A)} - f(p)|| + \alpha,0)\)
where:
The model yields performance advantages downstream user-based classification tasks.
We encourage usage and benchmarking on tasks involving:
Being exclusively trained on Reddit data, our models probably overfit to linguistic markers and traits which are relevant to characterizing the Reddit user population, but less salient in the general population. Domain-specific fine-tuning may be required before deployment.
Furthermore, our self-supervised approach enforces little or no control over biases, which models may actively use as part of their heuristics in contrastive and downstream tasks.