Downloads · 30 days
32
22% of all-time downloads
tmy100000001/LitDD_BERT
LitDD_BERT is a text classification model from tmy100000001. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
The abstract-screening classifier from LitDD, a pipeline that maps PubMed literature to gene–disease entries in the G2P developmental-disorder panel.
Downloads · 30 days
32
22% of all-time downloads
All-time downloads
148
Public
Parameters
396M
6.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
The abstract-screening classifier from LitDD, a pipeline that maps PubMed literature to gene–disease entries in the G2P developmental-disorder panel.
Given a title+abstract (TIAB), LitDD-BERT answers one question:
Does this paper present evidence that a gene causes a developmental disorder?
It is the first stage of a cascade. Its job is to reduce ~35M PubMed records to a tractable candidate set; precision is recovered downstream by a gene-mention gate (each abstract is restricted to the Gene2Phenotype entries of the genes it actually names) and an LLM adjudication step. See Intended use.
Current main is the add20k retrain (seed 44). The previously published checkpoint is
preserved at tag v1-seed42 and is not interchangeable with this one: it fires on
19.5% of random PubMed, this one on 0.51%. Both score ~0.92 F1 on the balanced test
set, which is nearly blind to the difference. If you are reproducing figures from the original
release, pin revision="v1-seed42".
The change is training-set composition, not architecture: 20,000 corpus-representative negatives — ordinary PubMed abstracts, decade-stratified — were added, lowering training prevalence from 51.8% to 24.0%. Every previous negative was drawn from gene–disease-relevant literature, so the model had never seen an ordinary chemistry, ecology, or clinical-trial abstract and had no reason to reject one.
transformers >= 4.48This is a ModernBERT architecture. Earlier versions fail with KeyError: 'modernbert'.
pip install "transformers>=4.48"
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "tmy100000001/LitDD_BERT"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()
tiab = ("A Novel Homozygous Nonsense HYDIN Gene Mutation p.(Arg951*) in Primary Ciliary "
"Dyskinesia. We report a consanguineous family in which two affected siblings ...")
enc = tok(tiab, truncation=True, max_length=8192, return_tensors="pt")
with torch.no_grad():
logits = model(**enc).logits
pred = logits.argmax(-1).item()
print(model.config.id2label[pred]) # RELEVANT / NOT_RELEVANT
Label mapping: 0 = NOT_RELEVANT, 1 = RELEVANT.
The deployed pipeline screens only English records published after 1980; matching that condition will reproduce the numbers below. ModernBERT's context is 8,192 tokens, which no PubMed abstract approaches (observed maximum ~800), so truncation is effectively never triggered.
Use it for: first-pass screening of biomedical abstracts for gene→developmental-disorder evidence, at corpus scale.
Do not use it for:
Recommended deployment: pair it with an upstream gene-mention gate. The pipeline admits an abstract only if a G2P panel gene is named in it (PubTator3 symbol annotations ∪ HGNC descriptive-name matching). The gate passes 81.7% of external curated truth papers but only 8.9% of random corpus papers, and composes with the screen: fire rate falls 0.51% → 0.28%.
| Base model | thomas-sounack/BioClinical-ModernBERT-large |
| Task | binary sequence classification, mean pooling, 2-way head |
| Training set | 37,335 (TIAB, label) pairs — 8,975 positive / 28,360 negative (24.0% positive) |
| Composition | clinician-annotated TIABs; molecular-mechanism-framed positives added after error analysis; literature from independent curated sources (gene-fold disciplined, no leakage into held-out genes); plus 20,000 corpus-representative negatives sampled from the full converted PubMed corpus, decade-stratified, excluding G2P-cited and curated-truth PMIDs |
| Hyperparameters | lr 3e-5, weight decay 0.1, 5 epochs, batch 32, seed 44 |
| Selection | 5-fold StratifiedGroupKFold on the training portion only; refit on full train; test touched once |
~11.7% of the annotated training records are titles with no abstract. These are kept deliberately: clinical-genetics titles conventionally state gene, variant and phenotype together, and title-only records are more likely to be positive (34.7%) than the corpus average (25.4%).
This checkpoint (seed 44) on the held-out annotated test set (n=2,779, 25.0% positive):
| Metric | Value |
|---|---|
| Precision | 0.898 |
| Recall | 0.938 |
| F1 | 0.918 |
| False-positive rate on a random PubMed sample (n=8,000) | 0.51% [0.38, 0.69] |
Across three training seeds (42/43/44): F1 0.9182 ± 0.0016, corpus fire rate
0.55 ± 0.05% (0.60 / 0.55 / 0.51). Note how much tighter the fire-rate spread is than in
the v1-seed42 release, where the same quantity ranged 5.00 / 5.85 / 10.37% — adding
corpus-representative negatives stabilised the property that a balanced test set cannot see.
Seed 44 is the median-F1 draw.
Measured on 1,752 papers from premined G2P / HPOA / ClinGen truth sets that are absent from the training set by normalised-TIAB match (not PMID — the training set carries no PMIDs, so a PMID join silently reports zero contamination):
| Denominator | Recall |
|---|---|
| All 1,752 papers | 0.763 [0.742, 0.782] |
| The 1,431 that pass the gene gate | 0.853 [0.834, 0.871] |
| The 637 "in-scope" papers (a G2P gene symbol/alias/full name appears in the TIAB) | 0.972 |
The whole-set figure is held down by paper types the model is trained to reject and that no downstream stage could curate anyway: reviews (n=77), pre-molecular linkage-only reports (n=98), and papers naming no gene in the abstract (n=700). Roughly 175 of the 1,752 were independently adjudicated as true negatives, so ~90% is the realistic ceiling on the raw denominator, not 100%.
Screen precision on a balanced test set does not transfer to PubMed. Measured directly on the papers this model plus the gene gate returns from a random corpus draw (all 248 fires among 87,600 corpus papers, manually adjudicated):
| Precision on returned papers | 0.565 [0.502, 0.625] |
| — at score ≥ 0.99 (80% of fires) | 0.648 |
| — at score 0.90–0.99 | 0.280 |
| — at score 0.58–0.90 | 0.167 |
This is a conservative lower bound: the corpus sample excluded every G2P-cited and curated-truth PMID, so already-curated true positives were removed by construction. Raising the decision threshold to 0.99 is a cheap precision gain if downstream volume matters.
Against EvAgg (gpt-4-turbo paper classifier), on identical rows:
| LitDD-BERT (this model) | EvAgg (gpt-4-turbo) | |
|---|---|---|
| Precision, annotated test set (n=2,739) | 0.892 | 0.534 |
| Recall, annotated test set | 0.937 | 0.995 |
| Precision on gated corpus fires | 0.708 [0.508, 0.851] | 0.191 [0.125, 0.283] |
| Fire rate, random PubMed (n≈8,000) | 0.51% | 3.13% [2.77, 3.53] |
The design trade is deliberate: substantially higher precision for slightly lower recall.
Comparison under an identical CV→refit→test protocol (same splits, same budget), measured on the original annotated test set and training recipe — a different operating point from the tables above, reported for transparency:
| Model | F1 |
|---|---|
| LitDD-BERT | 0.965 |
| BioClinical-ModernBERT-large | 0.925 |
| BioBERT | 0.923 |
| BiomedBERT | 0.923 |
| ModernBERT-large | 0.911 |
languages == "eng" and pubdate > 1980
before screening; the model was never trained or evaluated outside that slice.Manuscript under review; citation and DOI to follow. Code: https://github.com/ (repository link to be added at acceptance).
Please also cite the base model, BioClinical ModernBERT.