Downloads · 30 days
11
37% of all-time downloads
rucolaes/nih-adrd-pubmedbert
nih-adrd-pubmedbert is a text classification model from rucolaes. Use it when you need a label for a piece of text. The card lists the license as mit.
A binary text classifier that predicts whether a research project's title + abstract belongs to NIH's Alzheimer's Disease / Alzheimer's Disease-Related Dementias (AD/ADRD) research portfolio. Fine-tuned from PubMedBER…
Downloads · 30 days
11
37% of all-time downloads
All-time downloads
30
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
A binary text classifier that predicts whether a research project's title + abstract belongs to NIH's Alzheimer's Disease / Alzheimer's Disease-Related Dementias (AD/ADRD) research portfolio. Fine-tuned from PubMedBERT / BiomedBERT on NIH's own portfolio labels — not a general "neurodegenerative disease" classifier.
microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext (BERT-base,
domain vocabulary pretrained from scratch on PubMed abstracts + PMC full text)(title, abstract) as a tokenizer pair
(proper token_type_ids segmentation, not a manually concatenated string)Source: NIH ExPORTER bulk data (Projects + Abstracts,
joined on APPLICATION_ID), fiscal years 2018–2024.
Label: positive iff the award's NIH_SPENDING_CATS field contains one of these three
exact RCDC category strings (an explicit allowlist, not a substring heuristic):
Alzheimer's Disease (RCDC code 40)Alzheimer's Disease Related Dementias (ADRD) (code 3254)Alzheimer's Disease including Alzheimer's Disease Related Dementias (AD/ADRD) (code 3246)This is NIH's own strict scope — it does not include Parkinson's disease, ALS, or Huntington's disease, which NIH tracks under their own separate RCDC categories.
Preprocessing:
(title, abstract) pairs removed (renewal/resubmission text reused
verbatim across fiscal years — ~50% of the raw corpus)CORE_PROJECT_NUM (stratified, 70/15/15) so that no NIH project's
multi-year renewals cross the train/val/test boundary| Split | Rows | Positive | Rate |
|---|---|---|---|
| train | 125,961 | 11,451 | 9.1% |
| val | 42,431 | 2,411 | 5.7% |
| test | 42,034 | 2,513 | 6.0% |
Held-out test set (natural class distribution, never used for training or model selection):
| Metric | Value |
|---|---|
| PR-AUC | 0.960 |
| F1 (threshold 0.5) | 0.929 |
A near-duplicate leakage audit (TF-IDF cosine similarity between every test row and its nearest training neighbor) found PR-AUC stable (0.960 → 0.959) after excluding the most similar 1-2% of test rows, so this is not inflated by train/test text overlap.
On threshold choice: the raw softmax output is not a calibrated posterior probability — training uses negative downsampling plus inverse-frequency class weighting, which shifts the effective training prevalence well above NIH's true ~6% base rate. Use PR-AUC/ROC-AUC for ranking, and pick an operating threshold from a validation set matching your deployment distribution rather than assuming 0.5. Platt scaling (fit on the logit, not the raw probability) is a reasonable post-hoc calibration if you need an actual probability estimate.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("rucolaes/nih-adrd-pubmedbert")
model = AutoModelForSequenceClassification.from_pretrained("rucolaes/nih-adrd-pubmedbert").eval()
title = "..."
abstract = "..."
inputs = tokenizer(title, abstract, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
prob_adrd = torch.softmax(logits, dim=1)[0, 1].item()