Downloads · 30 days
25
5% of all-time downloads
HiTZ/Medical-mT5-xl-multitask
Medical-mT5-xl-multitask is a text generation model from HiTZ. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center" <br <img src="http://www.ixa.eus/sites/default/files/anitdote.png" style="width: 45%;" <h2 align="center"Medical mT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain</h2 <be
Downloads · 30 days
25
5% of all-time downloads
All-time downloads
457
Public
Parameters
3.7B
15 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors15 GB · 100%
From the Hugging Face model README
Medical MT5-xl-multitask is a version of Medical MT5 finetuned for sequence labelling. It can correctly label a wide range of Medical labels in unstructured text, such as Disease, Disability, ClinicalEntity, Chemical... Medical MT5-xl-multitask has been finetuned for English, Spanish, French and Italian, although it may work with a wide range of languages.
Medical MT5-xl-multitask was training using the Sequence-Labeling-LLMs library: https://github.com/ikergarcia1996/Sequence-Labeling-LLMs/
This library uses constrained decoding to ensure that the output contains the same words as the input and a valid HTML annotation. We recommend using Medical MT5-xl-multitask together with this library.
Although you can also directly use it with 🤗 huggingface. In order to label a sentence, you need to append the labels you wan to use, for example, if you want to label dieseases you should format your input as follows: <Disease> Torsade de pointes ventricular tachycardia during low dose intermittent dobutamine treatment in a patient with dilated cardiomyopathy and congestive heart failure .
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model = AutoModelForSeq2SeqLM.from_pretrained("Medical-MT5-xl-multitask",torch_dtype=torch.bfloat16, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("Medical-MT5-xl-multitask")
input_example = "<Disease> Torsade de pointes ventricular tachycardia during low dose intermittent dobutamine treatment in a patient with dilated cardiomyopathy and congestive heart failure ."
model_input = tokenizer(input_example, return_tensors="pt")
output = model.generate(**model_input.to(model.device),max_new_tokens=128,num_beams=1,num_return_sequences=1,do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
@misc{garcíaferrero2024medical,
title={Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain},
author={Iker García-Ferrero and Rodrigo Agerri and Aitziber Atutxa Salazar and Elena Cabrio and Iker de la Iglesia and Alberto Lavelli and Bernardo Magnini and Benjamin Molinet and Johana Ramirez-Romero and German Rigau and Jose Maria Villa-Gonzalez and Serena Villata and Andrea Zaninello},
year={2024},
eprint={2404.07613},
archivePrefix={arXiv},
primaryClass={cs.CL}
}