Downloads · 30 days
40
0% of all-time downloads
dennlinger/roberta-cls-consec
roberta-cls-consec is a text classification model from dennlinger. Use it when you need a label for a piece of text. It is set up for transformers.
This network has been fine-tuned for the task described in the paper Topical Change Detection in Documents via Embeddings of Long Sequences and is our best-performing base-transformer model. You can find more detailed…
Downloads · 30 days
40
0% of all-time downloads
All-time downloads
28.4K
Public
Parameters
125M
1.5 GB on disk
Likes
15
Public
Click a slice to open those files.
.bin501 MB · 33%
From the Hugging Face model README
This network has been fine-tuned for the task described in the paper Topical Change Detection in Documents via Embeddings of Long Sequences and is our best-performing base-transformer model. You can find more detailed information in our GitHub page for the paper here, or read the paper itself. The weights are based on RoBERTa-base.
The preferred way is through pipelines
from transformers import pipeline
pipe = pipeline("text-classification", model="dennlinger/roberta-cls-consec")
pipe("{First paragraph} [SEP] {Second paragraph}")
The model expects two segments that are separated with the [SEP] token. In our training setup, we had entire paragraphs as samples (or up to 512 tokens across two paragraphs), specifically trained on a Terms of Service data set. Note that this might lead to poor performance on "general" topics, such as news articles or Wikipedia.
The training task is to determine whether two text segments (paragraphs) belong to the same topical section or not. This can be utilized to create a topical segmentation of a document by consecutively predicting the "coherence" of two segments.
If you are experimenting via the Huggingface Model API, the following are interpretations of the LABELs:
LABEL_0: Two input segments separated by [SEP] do not belong to the same topic.LABEL_1: Two input segments separated by [SEP] do belong to the same topic.The results of this model can be found in the paper. We average over models from five different random seeds, which is why the specific results for this model might be different from the exact values in the paper.
Note that this model is not trained to work on classifying single texts, but only works with two (separated) inputs.