Downloads · 30 days
21
1% of all-time downloads
NotXia/pubmedbert-bio-ext-summ
pubmedbert-bio-ext-summ is a summarization model from NotXia. Use it when you need a shorter version of a longer text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
21
1% of all-time downloads
All-time downloads
1.6K
Public
Repo size
967 MB
Likes
1
Public
Click a slice to open those files.
.bin484 MB · 100%
From the Hugging Face model README
Work done for my Bachelor's thesis.
PubMedBERT fine-tuned
on MS^2 for extractive summarization.
The model architecture is similar to BERTSum.
Training code is available at biomed-ext-summ.
summarizer = pipeline("summarization",
model = "NotXia/pubmedbert-bio-ext-summ",
tokenizer = AutoTokenizer.from_pretrained("NotXia/pubmedbert-bio-ext-summ"),
trust_remote_code = True,
device = 0
)
sentences = ["sent1.", "sent2.", "sent3?"]
summarizer({"sentences": sentences}, strategy="count", strategy_args=2)
>>> (['sent1.', 'sent2.'], [0, 1])
Strategies to summarize the document:
length: summary with a maximum length (strategy_args is the maximum length).count: summary with the given number of sentences (strategy_args is the number of sentences).ratio: summary proportional to the length of the document (strategy_args is the ratio [0, 1]).threshold: summary only with sentences with a score higher than a given value (strategy_args is the minimum score).