Downloads · 30 days
21.1K
2% of all-time downloads
Angelakeke/RaTE-NER-Deberta
RaTE-NER-Deberta is a token classification model from Angelakeke. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of DeBERTa on the RaTE-NER dataset.
Downloads · 30 days
21.1K
2% of all-time downloads
All-time downloads
948K
Public
Parameters
184M
1.5 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors735 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of DeBERTa on the RaTE-NER dataset.
This model is trained to serve the RaTEScore metric, if you are interested in our pipeline, please refer to our paper and Github.
This model also can be used to extract Abnormality, Non-Abnormality, Anatomy, Disease, Non-Disease in medical radiology reports.
def run_ner(texts, idx2label, tokenizer, model, device): inputs = tokenizer(texts, max_length=512, padding=True, truncation=True, return_tensors="pt").to(device) with torch.no_grad(): outputs = model(**inputs) predicted_labels = torch.argmax(outputs.logits, dim=2).tolist() save_pairs = [] for i in range(len(texts)): predicted_entities = [idx2label[label] for label in predicted_labels[i]] non_pad_mask = inputs["input_ids"][i] != tokenizer.pad_token_id non_pad_length = non_pad_mask.sum().item() non_pad_input_ids = inputs["input_ids"][i][:non_pad_length] tokenized_text = tokenizer.convert_ids_to_tokens(non_pad_input_ids) save_pair = post_process(tokenized_text, predicted_entities, tokenizer) if i == 0: save_pairs = save_pair else: save_pairs.extend(save_pair) return save_pairs
ner_labels = ['B-ABNORMALITY', 'I-ABNORMALITY', 'B-NON-ABNORMALITY', 'I-NON-ABNORMALITY', 'B-DISEASE', 'I-DISEASE', 'B-NON-DISEASE', 'I-NON-DISEASE', 'B-ANATOMY', 'I-ANATOMY', 'O'] idx2label = {i: label for i, label in enumerate(ner_labels)}
tokenizer = AutoTokenizer.from_pretrained('Angelakeke/RaTE-NER-Deberta') model = AutoModelForTokenClassification.from_pretrained('Angelakeke/RaTE-NER-Deberta')
device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model.to(device) model.eval()
We recommend to inference by sentences.
text = ""
texts = text.split('. ') save_pair = run_ner(texts, idx2label, tokenizer, model, device)
</code></pre>
</details>Author: Weike Zhao
If you have any questions, please feel free to contact [email protected].
@inproceedings{zhao2024ratescore,
title={RaTEScore: A Metric for Radiology Report Generation},
author={Zhao, Weike and Wu, Chaoyi and Zhang, Xiaoman and Zhang, Ya and Wang, Yanfeng and Xie, Weidi},
booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing},
pages={15004--15019},
year={2024}
}