Downloads · 30 days
95
1% of all-time downloads
dnagpt/human_gpt2-v1
human_gpt2-v1 is a feature extraction model from dnagpt. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
dna language model trained using gpt2. using human genome data.
Downloads · 30 days
95
1% of all-time downloads
All-time downloads
8.4K
Public
Repo size
840 MB
Likes
10
Public
Click a slice to open those files.
.bin420 MB · 100%
From the Hugging Face model README
dna language model trained using gpt2. using human genome data.
Key features of our dangpt models:
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained('dnagpt/human_gpt2-v1')
tokenizer.tokenize("GAGCACATTCGCCTGCGTGCGCACTCACACACACGTTCAAAAAGAGTCCATTCGATTCTGGCAGTAG")
#result: [G','AGCAC','ATTCGCC',....]
model = AutoModel.from_pretrained('dnagpt/human_gpt2-v1')
import torch
dna = "ACGTAGCATCGGATCTATCTATCGACACTTGGTTATCGATCTACGAGCATCTCGTTAGC"
inputs = tokenizer(dna, return_tensors = 'pt')["input_ids"]
hidden_states = model(inputs)[0] # [1, sequence_length, 768]
# embedding with mean pooling
embedding_mean = torch.mean(hidden_states[0], dim=0)
print(embedding_mean.shape) # expect to be 768
# embedding with max pooling
embedding_max = torch.max(hidden_states[0], dim=0)[0]
print(embedding_max.shape) # expect to be 768