Downloads · 30 days
12
19% of all-time downloads
ptanwar/pubmedbert-chia-ner
pubmedbert-chia-ner is a token classification model from ptanwar. Use it when you need labels on individual words, such as names. The card lists the license as mit.
Fine-tuned microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext for named entity recognition on clinical trial eligibility criteria (CHIA corpus), as part of a course NLP project comparing fine-tuned biomedic…
Downloads · 30 days
12
19% of all-time downloads
All-time downloads
62
Public
Parameters
109M
436 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors436 MB · 100%
From the Hugging Face model README
Fine-tuned microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext for
named entity recognition on clinical trial eligibility criteria (CHIA corpus),
as part of a course NLP project comparing fine-tuned biomedical transformers
vs. GPT-4 prompting.
O + 15 entity types x B/I)| Learning rate | 5e-5 |
| Batch size | 8 |
| Epochs | 10 |
| Max sequence length | 256 |
| Adam epsilon | 1e-8 |
This checkpoint's own fold (entity-level, test set):
| Precision | Recall | F1 | |
|---|---|---|---|
| Strict (exact span match) | 0.639 | 0.674 | 0.656 |
| Relaxed (type + overlap match) | 0.750 | 0.791 | 0.770 |
Full 10-fold cross-validation (mean +/- std across all 10 folds -- the number directly comparable to Li et al. 2022's own 10-fold-averaged reporting; the weights hosted in this repo are from one of these 10 folds, not a checkpoint averaged across them):
| Precision | Recall | F1 | |
|---|---|---|---|
| Strict | 0.657 +/- 0.013 | 0.682 +/- 0.021 | 0.669 +/- 0.013 |
| Relaxed | 0.758 +/- 0.014 | 0.787 +/- 0.024 | 0.772 +/- 0.015 |
Both this checkpoint's own score and the 10-fold mean exceed Li et al. 2022's published PubMedBERT numbers on Chia (0.622 strict / 0.744 relaxed).
from transformers import AutoModelForTokenClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("ptanwar/pubmedbert-chia-ner")
model = AutoModelForTokenClassification.from_pretrained("ptanwar/pubmedbert-chia-ner")
Mood, Observation, Reference_point) -- see the full writeup for error analysis.