Downloads · 30 days
24
4% of all-time downloads
dnagpt/llama-dna
llama-dna is a machine learning model from dnagpt. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
We perform continuous pre-training of DNA sequence data based on the LLaMA model. This involves using a comprehensive and diverse dataset to further enhance the model's understanding and representation of genomic info…
Downloads · 30 days
24
4% of all-time downloads
All-time downloads
647
Public
Parameters
7B
14 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors14 GB · 100%
From the Hugging Face model README
We perform continuous pre-training of DNA sequence data based on the LLaMA model. This involves using a comprehensive and diverse dataset to further enhance the model's understanding and representation of genomic information. Specifically:
By continuously pre-training the LLaMA model with this DNA sequence data, we ensure that the model remains up-to-date with the latest genomic discoveries and maintains its ability to generalize well across different genomics tasks. This continuous learning process helps to improve the model's accuracy and robustness in handling complex biological sequences.
from transformers import AutoTokenizer, AutoConfig,AutoModel
from transformers import DataCollatorForLanguageModeling
from transformers import Trainer, TrainingArguments
from transformers import AutoConfig, AutoModelForCausalLM,LlamaForCausalLM,LlamaTokenizer
from tokenizers import Tokenizer
from datasets import load_dataset
tokenizer = LlamaTokenizer.from_pretrained("dnagpt/llama-dna")
tokenizer.pad_token = tokenizer.eos_token
model = LlamaForCausalLM.from_pretrained("dnagpt/llama-dna") #continue pretrain
text='''GCTGACTCTGCCAGGATGGAATGAAATTAGGTTGTTTTAATTATAATGTAAAGTCAGTTCTAGTCAGACATAGTCACATAGGCAAGTAAGGGAACCTAAAATTGCTTGGAAT,
The primary use of LLaMA is research on large language models, including'''
print(f"Tokenized by DNA-LLaMA tokenizer:{tokenizer.tokenize(text)}")
import torch
from transformers import pipeline
model_id = "dnagpt/llama-dna"
pipe = pipeline(
"text-generation",
model=model_id,
#torch_dtype=torch.bfloat16,
device_map="auto",
)
print(pipe("The key to life is"))
print(pipe("GGAATGAAATTAGGTTGTTTTAATTATAATGTAAAGTCAGTTCT"))