Downloads · 30 days
15
4% of all-time downloads
tgamstaetter/mult_tf
mult_tf is a text classification model from tgamstaetter. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
15
4% of all-time downloads
All-time downloads
407
Public
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.bin438 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of microsoft/BiomedNLP-PubMedBERT-base-uncased-abstract-fulltext on the None dataset. It achieves the following results on the evaluation set:
mult_tf is a fine-tuned [PubMedBERT](https://huggingface.co/microsoft/BiomedNLP-PubMedBERT-base
-uncased-abstract-fulltext)
model for 17-class medical specialty classification of biomedical text.
It distinguishes between 11 Internal Medicine sub-specialties and 6 other medical disciplines,
trained on 300,000 PubMed article titles using journal provenance as a distant supervision
signal. No manual annotation was used.
Companion to the binary classifier
tgamstaetter/im-bin-tf-abstr.
Fine-grained specialty classification of biomedical abstracts or titles
Not intended for: clinical decision support, diagnostic use, or any safety-critical application.
300,000 PubMed article titles from 77 medical journals. Labels from journal editorial scope
(distant supervision).
17 classes:
| Class label | Specialty |
|---|---|
angio | Angiology |
cardio | Cardiology |
endo | Endocrinology |
gastro | Gastroenterology |
geri | Geriatrics |
hemato | Hematology |
infect | Infectiology |
intens | Intensive Care Medicine |
nephro | Nephrology |
pulmo | Pulmonology |
rheu | Rheumatology |
anest | Anesthesiology |
gyn | Gynecology |
neuro | Neurology |
oto | Otorhinolaryngology |
psych | Psychiatry |
surgery | Surgery |
Dataset: Internal medicine and other specialties — Kaggle
Fine-tuned from microsoft/BiomedNLP-PubMedBERT-base-uncased-abstract-fulltext:
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 | Precision | Recall | Roc Auc |
|---|---|---|---|---|---|---|---|---|
| No log | 1.0 | 357 | 0.5694 | 0.8249 | 0.8243 | 0.8245 | 0.8249 | 0.9875 |
| 0.5397 | 2.0 | 714 | 0.5324 | 0.8324 | 0.8312 | 0.8313 | 0.8324 | 0.9890 |
| 0.523 | 3.0 | 1071 | 0.5193 | 0.8354 | 0.8348 | 0.8346 | 0.8354 | 0.9895 |
| 0.523 | 4.0 | 1428 | 0.5180 | 0.8364 | 0.8358 | 0.8355 | 0.8364 | 0.9896 |
Evaluated on a held-out test set of 100,000 titles:
| Metric | Value |
|---|---|
| Accuracy | 0.835 |
| Macro F1 | 0.834 |
| Macro Precision | 0.836 |
| Macro Recall | 0.835 |
| ROC-AUC (macro OvR) | 0.903 |
Lowest per-class F1 scores: intensive care medicine (0.670), geriatrics (0.683),
angiology (0.704) — reflecting known clinical content overlap with adjacent specialties.
Transformers 4.31.0
Pytorch 2.0.1+cu118
Datasets 2.13.1
Tokenizers 0.13.3
@misc{gamstaetter2023modelmc,
author = {Gamstaetter, Thomas},
title = {mult\_tf: Fine-tuned {PubMedBERT} for multiclass medical specialty
classification},
year = {2023},
howpublished = {Hugging Face},
url = {https://huggingface.co/tgamstaetter/mult_tf}
}
Associated preregistration: OSF — DOI 10.17605/OSF.IO/XFDBV