Downloads · 30 days
3.6K
3% of all-time downloads
COGNANO/VHHBERT
VHHBERT is a fill-mask model from COGNANO. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
VHHBERT is a RoBERTa-based model pre-trained on two million VHH sequences in VHHCorpus-2M. VHHBERT has the same model parameters as RoBERTa<subBASE</sub, except that it used positional embeddings with a length of 185…
Downloads · 30 days
3.6K
3% of all-time downloads
All-time downloads
133K
Public
Parameters
85.8M
343 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors343 MB · 100%
From the Hugging Face model README
VHHBERT is a RoBERTa-based model pre-trained on two million VHH sequences in VHHCorpus-2M. VHHBERT has the same model parameters as RoBERTa<sub>BASE</sub>, except that it used positional embeddings with a length of 185 to cover the maximum sequence length of 179 in VHHCorpus-2M. Further details on VHHBERT are described in our paper "A SARS-CoV-2 Interaction Dataset and VHH Sequence Corpus for Antibody Language Models.”
The model and tokenizer can be loaded using the transformers library.
from transformers import BertTokenizer, RobertaModel
tokenizer = BertTokenizer.from_pretrained("COGNANO/VHHBERT")
model = RobertaModel.from_pretrained("COGNANO/VHHBERT")
If you use VHHBERT in your research, please cite the following paper.
@inproceedings{tsuruta2024sars,
title={A {SARS}-{C}o{V}-2 Interaction Dataset and {VHH} Sequence Corpus for Antibody Language Models},
author={Hirofumi Tsuruta and Hiroyuki Yamazaki and Ryota Maeda and Ryotaro Tamura and Akihiro Imura},
booktitle={Advances in Neural Information Processing Systems 37},
year={2024}
}