Downloads · 30 days
41
14% of all-time downloads
jackleejm/distilbert-medication-ner
distilbert-medication-ner is a token classification model from jackleejm. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of distilbert-base-cased on synthetically generated medication data by Synthea.
Downloads · 30 days
41
14% of all-time downloads
All-time downloads
302
Public
Parameters
65.2M
261 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors261 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of distilbert-base-cased on synthetically generated medication data by Synthea.
More details on how this model was trained can be found on GitHub.
A fine-tuned NER model developed to handle 5 specific entities (i.e. DRUG, DOSAGE, ROUTE, BRAND, QUANTITY) when processing medication strings such as:
The model was trained and evaluated on limited manually annotated datasets (i.e. train_n_samples=309, eval_n_samples=335), achieved the following evaluation metrics:
from transformers import AutoTokenizer, AutoModelForTokenClassification
model_name = "jackleejm/distilbert-medication-ner"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
from transformers import pipeline
ner_pipeline = pipeline(
task="token-classification",
model=model,
tokenizer=tokenizer,
aggregation_strategy="simple",
device_map="auto",
)
input = ["Acetaminophen 325 MG Oral Tablet"]
results = ner_pipeline(input)
print(results)
# Outputs
[
[
{
"word": "Acetaminophen",
"score": np.float32(0.99948627),
"entity_group": "DRUG",
"start": 0,
"end": 13
},
{
"word": "325 MG",
"score": np.float32(0.99882394),
"entity_group": "DOSAGE",
"start": 14,
"end": 20
},
{
"word": "Oral Tablet",
"score": np.float32(0.9994621),
"entity_group": "ROUTE",
"start": 21,
"end": 32
}
]
]