Downloads · 30 days
31
13% of all-time downloads
vojtam/DNAGPT2_16
DNAGPT2_16 is a text generation model from vojtam. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
DNAGPT2 is a family of autoregressive (decoder-only) transformer models trained on genomic DNA sequences.
Downloads · 30 days
31
13% of all-time downloads
All-time downloads
235
Public
Parameters
85.9M
343 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors343 MB · 100%
From the Hugging Face model README
DNAGPT2 is a family of autoregressive (decoder-only) transformer models trained on genomic DNA sequences.
The models follow the GPT-2 architecture and are trained from scratch on a multi-species genome dataset.
These models are designed for:
Input: Raw DNA sequences containing the characters A, C, G, T.
Output: Logits/Probabilities for the next token in the sequence.
The models were pretrained on the dataset provided by the authors of DNABERT-2.
The models were trained using the PyTorch framework and the nanoGPT recipe.
The models were evaluated on their ability to compress DNA sequences (measured in bits per symbol or bps) using Arithmetic Encoding. Lower is better.
| Dataset | Metric | DNAGPT2_32 | Benchmark (gzip -9) | Benchmark (Jarvis3) |
|---|---|---|---|---|
| Homo sapiens (T2T-CHM13v2.0) | bits/symbol | 1.470 | 2.022 | 1.384 |
| M. llanfair... (Bacteria) | bits/symbol | 1.783 | 2.142 | 1.713 |
| A. thaliana (Plant - Chr1) | bits/symbol | 1.876 | 2.161 | 1.702 |
The DNAGPT2_32 model outperforms general-purpose compressors (gzip) and competitive deep learning models like hyenaDNA and megaDNA on the evaluated datasets.
The model is compatible with the Hugging Face transformers library.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
# Select the model variant (e.g., vocab size 128 or 32)
# Replace with the specific repository path if hosted on HF Hub
hf_model_repository = "vojtam/DNAGPT2_128"
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
hf_model_repository,
trust_remote_code=True
).to(device)
tokenizer = AutoTokenizer.from_pretrained(
hf_model_repository,
trust_remote_code=True
)
# Inference Example
dna_sequence = "ACGTTGCAAACG"
token_ids = tokenizer.encode(dna_sequence, return_tensors="pt").to(device)
with torch.no_grad():
logits = model(token_ids).logits
print(f"Input: {dna_sequence}")
print(f"Logits shape: {logits.shape}")