Downloads · 30 days
15
23% of all-time downloads
Siddharth63/medul2-base
medul2-base is a machine learning model from Siddharth63. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
PubMedUL2 and MedUL2 are a family of domain-specific UL2/T5-style encoder–decoder language models pretrained on large-scale biomedical and medical corpora using the UL2 (Mixture-of-Denoisers) objective.
Downloads · 30 days
15
23% of all-time downloads
All-time downloads
66
Public
Parameters
248M
2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors990 MB · 50%
From the Hugging Face model README
PubMedUL2 and MedUL2 are a family of domain-specific UL2/T5-style encoder–decoder language models pretrained on large-scale biomedical and medical corpora using the UL2 (Mixture-of-Denoisers) objective.
These checkpoints are pretraining-only models and must be fine-tuned before use on downstream tasks.
These models were pretrained using UL2, a unified framework that formulates language modeling objectives as denoising tasks.
UL2 introduces a Mixture-of-Denoisers (MoD) approach that samples from multiple denoising paradigms during pretraining.
UL2 pretraining uses a mixture of three denoising tasks:
R-denoising (Regular Span Corruption)
X-denoising (Extreme Span Corruption)
S-denoising (Sequential / PrefixLM)
During pretraining, a paradigm token is inserted at the beginning of each input:
| Token | Mode | Recommended Use |
|---|---|---|
[NLU] | R-denoising | Classification, QA, retrieval |
[NLG] | X-denoising | Mixed understanding & generation |
[S2S] | S-denoising | Generative / causal tasks |
Important:
For best performance, the same token should be prepended during fine-tuning and inference.
T5ForConditionalGenerationThese models are intended to be fine-tuned for:
These checkpoints are self-supervised pretraining models only and require task-specific fine-tuning.
[NLU], [NLG], or [S2S] to input text[NLU][S2S][NLG]| Model Name | Parameter Count | Description | Access |
|---|---|---|---|
pubmedul2-tiny-nl6 | 19.26M | Tiny UL2-style model with 6 layers | Open |
pubmedul2-mini-nl8 | 50.12M | Mini UL2 with 8 layers | Open |
pubmedul2-small | 60.52M | Small UL2 variant | Open |
pubmedul2-small-nl24 | 192.73M | Small UL2 with 24 layers | Open |
medul2-base | 222.93M | Base UL2/T5-style model | Open |
pubmedul2-base | 222.93M | Base UL2/T5-style model | Open |
medul2-base-nl36 | 619.44M | Base UL2 with 36 layers | Gated commercial |
pubmedul2-base-nl36 | 619.44M | Base UL2 with 36 layers | Gated commercial |
medul2-large | 737.72M | Large UL2/T5-style model | Gated non-commercial |
pubmedul2-large | 737.72M | Large UL2/T5-style model | Gated non-commercial |
medul2-large-nl36 | 1090.14M | Very large UL2 with 36 layers | Access on Request |
We evaluate PubMedUL2 and MedUL2 models on a biomedical Named Entity Recognition (NER) task using multiple matching criteria to better capture boundary-level performance.
The evaluation reports entity-level F1 scores across different biomedical entity types and model sizes.
An entity prediction is considered correct only if both the entity span and label exactly match the gold annotation.
| entity_type | medul2-base | pubmedul2-base | pubmedul2-mini-nl8 | pubmedul2-small | pubmedul2-tiny-nl6 |
|---|---|---|---|---|---|
| cell_line | 0.42 | 0.43 | 0.44 | 0.43 | 0.35 |
| cell_type | 0.59 | 0.58 | 0.59 | 0.58 | 0.52 |
| chemical | 0.76 | 0.75 | 0.72 | 0.72 | 0.56 |
| disease | 0.7 | 0.73 | 0.7 | 0.68 | 0.63 |
| dna | 0.59 | 0.55 | 0.54 | 0.55 | 0.45 |
| gene | 0.62 | 0.59 | 0.6 | 0.59 | 0.55 |
| protein | 0.59 | 0.58 | 0.58 | 0.59 | 0.55 |
| rna | 0.6 | 0.56 | 0.55 | 0.6 | 0.56 |
| species | 0.66 | 0.67 | 0.58 | 0.63 | 0.54 |
A prediction is counted as correct if it partially overlaps with a gold entity of the same type.
| entity_type | medul2-base | pubmedul2-base | pubmedul2-mini-nl8 | pubmedul2-small | pubmedul2-tiny-nl6 |
|---|---|---|---|---|---|
| cell_line | 0.48 | 0.49 | 0.48 | 0.48 | 0.41 |
| cell_type | 0.66 | 0.64 | 0.66 | 0.65 | 0.59 |
| chemical | 0.79 | 0.78 | 0.76 | 0.75 | 0.6 |
| disease | 0.82 | 0.84 | 0.8 | 0.79 | 0.74 |
| dna | 0.65 | 0.61 | 0.6 | 0.61 | 0.53 |
| gene | 0.76 | 0.74 | 0.74 | 0.73 | 0.68 |
| protein | 0.66 | 0.66 | 0.66 | 0.67 | 0.64 |
| rna | 0.68 | 0.63 | 0.64 | 0.66 | 0.65 |
| species | 0.68 | 0.7 | 0.61 | 0.65 | 0.56 |
Predictions are evaluated using Intersection-over-Union (IoU) overlap between predicted and gold spans, providing a softer boundary-based metric.
| entity_type | medul2-base | pubmedul2-base | pubmedul2-mini-nl8 | pubmedul2-small | pubmedul2-tiny-nl6 |
|---|---|---|---|---|---|
| cell_line | 0.5 | 0.5 | 0.5 | 0.5 | 0.42 |
| cell_type | 0.67 | 0.66 | 0.68 | 0.67 | 0.62 |
| chemical | 0.83 | 0.83 | 0.82 | 0.82 | 0.72 |
| disease | 0.85 | 0.86 | 0.86 | 0.85 | 0.82 |
| dna | 0.65 | 0.62 | 0.62 | 0.62 | 0.55 |
| gene | 0.76 | 0.75 | 0.75 | 0.74 | 0.71 |
| protein | 0.67 | 0.66 | 0.67 | 0.67 | 0.66 |
| rna | 0.68 | 0.65 | 0.66 | 0.67 | 0.67 |
| species | 0.72 | 0.74 | 0.65 | 0.69 | 0.58 |
This project would not have been possible without compute generously provided by Google TPU Research Cloud.
Thanks to:
Please refer to the individual model repositories for license and access details, which may vary depending on training data sources.