Downloads · 30 days
222
9% of all-time downloads
AventIQ-AI/bert-medical-entity-extraction
bert-medical-entity-extraction is a machine learning model from AventIQ-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository hosts the quantized version of the bert-base-cased model for Medical Entity Extraction using the 'tner/bc5cdr' dataset. The model is specifically designed to recognize entities related to Disease,Sympt…
Downloads · 30 days
222
9% of all-time downloads
All-time downloads
2.5K
Public
Parameters
108M
215 MB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors215 MB · 100%
From the Hugging Face model README
This repository hosts the quantized version of the bert-base-cased model for Medical Entity Extraction using the 'tner/bc5cdr' dataset. The model is specifically designed to recognize entities related to Disease,Symptoms,Drug. The model has been optimized for efficient deployment while maintaining high accuracy, making it suitable for resource-constrained environments.
tner/bc5cdrpip install transformers torch
from transformers import BertTokenizerFast, BertForTokenClassification
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
model_name = "AventIQ-AI/bert-medical-entity-extraction"
model = BertForTokenClassification.from_pretrained(model_name).to(device)
tokenizer = BertTokenizerFast.from_pretrained(model_name)
from transformers import pipeline
ner_pipeline = pipeline("ner", model=model_name, tokenizer=tokenizer)
test_sentence = "An overdose of Ibuprofen can lead to severe gastric issues."
ner_results = ner_pipeline(test_sentence)
label_map = {
"LABEL_0": "O", # Outside (not an entity)
"LABEL_1": "Drug",
"LABEL_2": "Disease",
"LABEL_3": "Symptom",
"LABEL_4": "Treatment"
}
def merge_tokens(ner_results):
merged_entities = []
current_word = ""
current_label = ""
current_score = 0
count = 0
for entity in ner_results:
word = entity["word"]
label = entity["entity"] # Model's output (e.g., LABEL_1, LABEL_2)
score = entity["score"]
# Merge subwords
if word.startswith("##"):
current_word += word[2:] # Remove '##' and append
current_score += score
count += 1
else:
if current_word: # Store the previous merged word
mapped_label = label_map.get(current_label, "Unknown")
merged_entities.append((current_word, mapped_label, current_score / count))
current_word = word
current_label = label
current_score = score
count = 1
# Add the last word
if current_word:
mapped_label = label_map.get(current_label, "Unknown")
merged_entities.append((current_word, mapped_label, current_score / count))
return merged_entities
print("\n🩺 Medical NER Predictions:")
for word, label, score in merge_tokens(ner_results):
if label != "O": # Skip non-entities
print(f"🔹 Entity: {word} | Category: {label} | Score: {score:.4f}")
| Entity Type | Precision | Recall | F1 Score | Number of Entities |
|---|---|---|---|---|
| Disease | 91.46% | 92.07% | 91.76% | 3,000 |
| Drug | 71.25% | 72.83% | 72.03% | 1,266 |
| Symptom | 89.83% | 93.02% | 91.40% | 3,524 |
| Treatment | 88.83% | 92.02% | 90.40% | 3,124 |
The Hugging Face's tner/bc5cdr dataset was used, containing texts and their ner tags.
Post-training quantization was applied using PyTorch's built-in quantization framework to reduce the model size and improve inference efficiency.
.
├── model/ # Contains the quantized model files
├── tokenizer_config/ # Tokenizer configuration and vocabulary files
├── model.safetensors/ # Quantized Model
├── README.md # Model documentation
Contributions are welcome! Feel free to open an issue or submit a pull request if you have suggestions or improvements.