Downloads · 30 days
62
11% of all-time downloads
brineylab/preferential-250k
preferential-250k is a fill-mask model from brineylab. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
Preferential-250k is an antibody language model that uses an ESM-2 architecture. It was pre-trained on paired sequences from Jaffe et al. and Hurtado et al. Datasets used for pre-training are available on Zenodo and c…
Downloads · 30 days
62
11% of all-time downloads
All-time downloads
562
Public
Parameters
356M
1.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 100%
From the Hugging Face model README
Preferential-250k is an antibody language model that uses an ESM-2 architecture. It was pre-trained on paired sequences from Jaffe et al. and Hurtado et al. Datasets used for pre-training are available on Zenodo and code is available on GitHub. More details can be found in our paper published in Patterns.
Load the model and tokenizer as follows:
from transformers import EsmTokenizer, EsmForMaskedLM
model = EsmForMaskedLM.from_pretrained("brineylab/preferential-250k")
tokenizer = EsmTokenizer.from_pretrained("brineylab/preferential-250k")
The tokenizer expects sequences formatted as: HEAVY_CHAIN<cls><cls>LIGHT_CHAIN.
The model can be finetuned for classification tasks (such as specificity and pair classification in the paper) by loading the model with a sequence classification head:
from transformers import EsmForSequenceClassification
model = EsmForSequenceClassification.from_pretrained("brineylab/preferential-250k")
# freeze the base model weights prior to finetuning
for param in model.base_model.parameters():
param.requires_grad = False