Downloads · 30 days
47
16% of all-time downloads
hu-lab/PlantGFM
PlantGFM is a machine learning model from hu-lab. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
PlantGFM is a genetic foundation model pre-trained on the complete genome sequences of 12 model plants, encompassing 108 billion nucleotides. Using the Hyena framework with 220 million parameters and a context length…
Downloads · 30 days
47
16% of all-time downloads
All-time downloads
288
Public
Repo size
1.9 GB
Likes
1
Public
Click a slice to open those files.
.bin941 MB · 100%
From the Hugging Face model README
PlantGFM is a genetic foundation model pre-trained on the complete genome sequences of 12 model plants, encompassing 108 billion nucleotides. Using the Hyena framework with 220 million parameters and a context length of 64K bp, PlantGFM models sequences at single-nucleotide resolution. The model employed a length warm-up strategy, starting with 1K bp fragments and gradually increasing to 64K bp, enhancing training stability and accelerating convergence.
Developed by: hu-lab
Install the runtime library first:
pip install transformers
To calculate the embedding of a dna sequence:
import torch
from transformers import PreTrainedTokenizerFast
from plantgfm.modeling_plantgfm import PlantGFMForCausalLM
from plantgfm.configuration_plantgfm import PlantGFMConfig
config = PlantGFMConfig.from_pretrained("hu-lab/PlantGFM")
tokenizer = PreTrainedTokenizerFast.from_pretrained("hu-lab/PlantGFM")
model = PlantGFMForCausalLM.from_pretrained("hu-lab/PlantGFM", config=config)
sequences = ["CCCTAAACCCTAAACCCTAAA", "ATGGCGTGGCTG"]
# get single-nucleotide sequences with space between each base
single_nucleotide_sequences = list(map(lambda seq: " ".join(list(seq)), sequences))
tokenized_sequences = tokenizer(single_nucleotide_sequences, padding="longest")["input_ids"]
input_ids = torch.LongTensor(tokenized_sequences)
embd = model(input_ids=input_ids, output_hidden_states=True)["hidden_states"][0]
print(embd)
Model was trained for 468 hours on 8 Nvidia A800-80G GPUs.