Downloads · 30 days
9
16% of all-time downloads
keiwoo/Pep2Tcr-TCRGen
Pep2Tcr-TCRGen is a text generation model from keiwoo. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This repository contains relevant information of our research article 'Leveraging contrastive learning for conditional and robust TCR sequence generation'.
Downloads · 30 days
9
16% of all-time downloads
All-time downloads
57
Public
Parameters
179M
715 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors715 MB · 100%
From the Hugging Face model README
This repository contains relevant information of our research article 'Leveraging contrastive learning for conditional and robust TCR sequence generation'.
You can use Pep2Tcr trained on TCRGen dataset with Transformers. To get started, install the necessary dependencies to setup your environment:
pip install -U transformers torch
from transformers import EncoderDecoderModel, AutoTokenizer
import torch
torch.cuda.manual_seed_all(42)
device = torch.device('cuda:1')
model_name = 'keiwoo/Pep2Tcr-TCRGen'
model = EncoderDecoderModel.from_pretrained(
model_name,
dtype=torch.float16, # If your device does not support `BF16` precision mode
# (GPU based on NVIDIA's Ampere and its subsequent architectures), you should comment this line.
).to(device)
model.eval()
tokenizer = AutoTokenizer.from_pretrained(model_name, do_lower_case=False)
sample_kwargs = {
'do_sample': True,
'top_k': 9,
'top_p': 0.93,
'temperature': 0.9,
'num_return_sequences': 16,
'repetition_penalty': 1.1,
'no_repeat_ngram_size': 3,
'max_new_tokens': 24,
}
beam_kwargs = {
'num_beams': 16,
'num_return_sequences': 16,
'repetition_penalty': 1.1,
'no_repeat_ngram_size': 3,
'max_new_tokens': 24,
}
inputs = "AAGIGILTV"
tokenized = tokenizer(" ".join(inputs), return_tensors="pt", ).to(device)
outputs = model.generate(**tokenized, **sample_kwargs) # If you want different sequences every generation, uncomment this line.
# outputs = model.generate(**tokenized, **beam_kwargs) # If you want same sequences every generation (most possible sequence the model think), uncomment this line.
tcr = tokenizer.batch_decode(outputs, skip_special_tokens=True)
tcr = [i.replace(' ', '') for i in tcr]
print(tcr)
# ['CASSQEGLAGGTDTQYF', 'CASSQEGLAGGYNEQFF', 'CASSQDPGLAGGYNEQFF', ...]
Please refer to https://github.com/keiwoo/Pep2Tcr