Downloads · 30 days
38
100% of all-time downloads
ItsnotAilabs/AlphaGenome-50M
AlphaGenome-50M is a feature extraction model from ItsnotAilabs. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
38
100% of all-time downloads
All-time downloads
38
Public
Parameters
64.5M
417 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors258 MB · 99%
From the Hugging Face model README
AlphaGenome-50M is a state-of-the-art genomic sequence embedding and variant impact prediction model. Fine-tuned from the InstaDeepAI/nucleotide-transformer-v2-50m-multi-species base model, it is specifically designed for non-coding variant analysis and regulatory element classification.
The model leverages the Nucleotide Transformer architecture to capture complex genomic patterns. It uniquely provides Variant Impact Scores (AVI) and facilitates single-variant effect analysis across RNA-seq, DNASE, and ChIP assays. Furthermore, it supports tissue-type ontology resolution (UBERON/CL), saturation mutagenesis window scanning, and seamless extraction of GENCODE v46 coordinates.
The primary use cases for AlphaGenome-50M include:
Sequences are represented using 6-mer tokenization with a BPE nucleotide vocabulary. The max context length is 1024 tokens. Input strings should be pure DNA sequences consisting of A, T, C, G characters.
AlphaGenome-50M is based on the Nucleotide Transformer v2 architecture:
While AlphaGenome-50M is a sequence feature extraction model rather than an instruction-tuned LLM, it requires specific input formatting. Do not use special control tokens like <s> or [CLS] manually unless bypassing the tokenizer. Pass the raw nucleotide string.
Sequence format: <DNA_STRING>
Example: ATGCGTACGTTAGCTAGCTAGCTAGCTAGCTAGC
| Format / Precision | Memory Footprint | Latency (CPU, per seq) | Latency (GPU T4, per seq) |
|---|---|---|---|
| FP32 (Base) | ~200 MB | 45 ms | 12 ms |
| FP16 (Half) | ~100 MB | 30 ms | 6 ms |
| INT8 (Quantized) | ~50 MB | 20 ms | 4 ms |
| GGUF/Q4_K_M | ~30 MB | 15 ms | N/A |
AlphaGenome-50M was evaluated on several genomic benchmarks:
| Benchmark / Task | Metric | Score (Estimated) |
|---|---|---|
| ClinVar Pathogenic | AUC | 0.87 |
| ENCODE cCRE Classification | F1 | 0.82 |
| Variant Effect Correlation | Spearman | 0.71 |
| Empirical Genomic Variant Impact | AUC | 0.4810 |
| Sequence Processing Throughput | seq/s | 8.93 |
You can use this model with the transformers library to encode DNA sequences and integrate with Biopython to parse FASTA files:
from transformers import AutoTokenizer, AutoModel
from Bio import SeqIO
import torch
# Load model and tokenizer
model_name = "MedinaMemorySystems/AlphaGenome-50M"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
# Parse a sequence using Biopython
record = next(SeqIO.parse("example.fasta", "fasta"))
dna_sequence = str(record.seq)
# Encode the DNA sequence
inputs = tokenizer(dna_sequence, return_tensors="pt", truncation=True, max_length=1024)
with torch.no_grad():
outputs = model(**inputs)
# Extract sequence embeddings
embeddings = outputs.last_hidden_state
print(f"Embedding shape: {embeddings.shape}")
@misc{medinamemorysystems2026alphagenome,
title={AlphaGenome-50M: Genomic Sequence Embedding and Variant Impact Prediction},
author={MedinaMemorySystems},
year={2026},
publisher={Hugging Face}
}