Downloads · 30 days
24
100% of all-time downloads
ItsnotAilabs/AlphaFold-Embed-8M
AlphaFold-Embed-8M is a feature extraction model from ItsnotAilabs. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
24
100% of all-time downloads
All-time downloads
24
Public
Parameters
7.4M
29.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors29.5 MB · 100%
From the Hugging Face model README
AlphaFold-Embed-8M is a compact yet highly effective protein sequence embedding model designed for structural confidence prediction, domain boundary analysis, and disorder assessment. Fine-tuned from facebook/esm2_t6_8M_UR50D, this model bridges sequence and structure by learning to predict structural features directly from primary amino acid sequences.
It uniquely provides per-residue pLDDT (predicted local distance difference test) confidence prediction, domain boundary detection, and intrinsic disorder scoring. Furthermore, it supports UniProt ID resolution and generates structural homology embeddings that are compatible with Foldseek.
The primary use cases for AlphaFold-Embed-8M include:
Sequences are represented directly as amino acid characters (e.g., "MKTII..."). The model uses standard ESM tokenization. The maximum sequence length is 1024 residues; longer proteins should be truncated or processed in overlapping chunks.
AlphaFold-Embed-8M leverages the ESM-2 architecture, optimized for efficiency:
AlphaFold-Embed-8M takes pure amino acid sequences. No special prompting template is required, but input strings must contain only valid amino acid character tokens.
Sequence format: <AMINO_ACID_STRING>
Example: MKTIIALSYIFCLVFADYKDDDDK
| Format / Precision | Memory Footprint | Latency (CPU, per seq) | Latency (GPU T4, per seq) |
|---|---|---|---|
| FP32 (Base) | ~32 MB | 15 ms | 5 ms |
| FP16 (Half) | ~16 MB | 10 ms | 3 ms |
| INT8 (Quantized) | ~8 MB | 8 ms | 2 ms |
| GGUF/Q4_K_M | ~5 MB | 6 ms | N/A |
AlphaFold-Embed-8M was evaluated on multiple structural biology benchmarks:
| Benchmark / Task | Metric | Score (Estimated) |
|---|---|---|
| pLDDT Prediction | MAE | 5.2 |
| Domain Boundary Detection | F1 | 0.78 |
| Intrinsic Disorder Scoring | AUC | 0.91 |
| CASP15 GDT-TS | Pearson Correlation | 0.83 |
| Empirical pLDDT MAE | MAE | 24.9567 |
| Sequence Processing Throughput | seq/s | 5.01 |
You can use this model with the transformers library to extract per-residue embeddings and predict structural properties, seamlessly integrating with Biopython:
from transformers import AutoTokenizer, AutoModel
from Bio import SeqIO
import torch
# Load model and tokenizer
model_name = "MedinaMemorySystems/AlphaFold-Embed-8M"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
# Parse a sequence using Biopython
record = next(SeqIO.parse("protein.fasta", "fasta"))
protein_sequence = str(record.seq)
# Encode the protein sequence
inputs = tokenizer(protein_sequence, return_tensors="pt", truncation=True, max_length=1024)
with torch.no_grad():
outputs = model(**inputs)
# Extract per-residue embeddings
embeddings = outputs.last_hidden_state
print(f"Embedding shape: {embeddings.shape}")
@misc{medinamemorysystems2026alphafoldembed,
title={AlphaFold-Embed-8M: Efficient Structural Feature Prediction from Protein Sequences},
author={MedinaMemorySystems},
year={2026},
publisher={Hugging Face}
}