Downloads · 30 days
46
100% of all-time downloads
devansh0703/PeptEdgeV2
PeptEdgeV2 is a machine learning model from devansh0703. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for custom. The card lists the license as mit.
State-of-the-art antimicrobial peptide (AMP) classifier with 189x fewer parameters than prior methods, reaching 91.47% F1, 91.34% accuracy, and 95.10% AUC on the GenPept-Curated-2025 benchmark.
Downloads · 30 days
46
100% of all-time downloads
All-time downloads
46
Public
Parameters
3.4M
13.8 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors13.8 MB · 100%
From the Hugging Face model README
State-of-the-art antimicrobial peptide (AMP) classifier with 189x fewer parameters than prior methods, reaching 91.47% F1, 91.34% accuracy, and 95.10% AUC on the GenPept-Curated-2025 benchmark.
| Metric | PeptEdgeV2 | ESM-2 LoRA (prior SOTA) | Delta |
|---|---|---|---|
| F1 Score | 91.47% | 88.30% | +3.17% |
| Accuracy | 91.34% | 86.80% | +4.54% |
| AUC | 0.9510 | 0.938 | +0.013 |
| Parameters | 3.43M | 650M | 189x fewer |
| Inference | 52,800 seq/s | 248 seq/s | 213x faster |
PeptEdgeV2 combines four components in a single compact encoder:
Conv1D branches (kernel sizes 3, 5, 7, 11) with
batch norm + residual connections.Total parameters: 3,434,018.
config.json): d_model=192, n_heads=6, num_layers=5, ff_dim=384, dropout=0.25, sd_prob=0.05, vocab_size=21, max_len=200, num_classes=20..20 (PAD=0, 20 standard AAs, X=20), truncated/padded
to max_len=200.Install:
pip install torch huggingface-hub safetensors numpy pandas scikit-learn
The repo contains full source (model_v2.py, data_utils.py, train_final.py). To load the model:
import torch
from model_v2 import PeptEdgeV2
model = PeptEdgeV2.from_pretrained("devansh0703/PeptEdgeV2")
model.eval()
# Encode a peptide (tokens 1-20 = AAs, 0 = pad)
from data_utils import encode_sequence
ids = encode_sequence("GLFDIVKKVVGALGSL", max_len=200)
x = torch.tensor([ids])
with torch.no_grad():
logits = model(x)
prob_amp = torch.softmax(logits, dim=-1)[0, 1].item()
print(f"P(AMP) = {prob_amp:.3f}")
Trained on the GenPept-Curated-2025 benchmark. The curated tables are mirrored on the Hub as a dataset repo:
Dataset summary:
Note on splits: the published numbers come from a simple random stratified split (
random_state=42), not the homology-controlled CD-HIT split described in the README/paper.
The final pipeline lives in train_final.py (CUDA + mixed precision required):
python train_final.py
Key hyperparameters (embedded in train_final.py): lr=3e-4, weight_decay=5e-5,
label_smoothing=0.1, warmup 15 epochs over 100, cosine decay, gradient clipping 1.0,
early stopping patience 30 on validation F1.
From results/final_results.json:
| Metric | Value |
|---|---|
| Accuracy | 0.9134 |
| Precision | 0.9009 |
| Recall | 0.9290 |
| F1 | 0.9147 |
| AUC | 0.9510 |
| MCC | 0.8272 |
| Params | 3,434,018 |
| File | Description |
|---|---|
config.json | Model hyperparameters |
model.safetensors | Trained weights (SOTA pipeline) |
model_v2.py | PeptEdgeV2 architecture + from_pretrained/save_pretrained |
data_utils.py | AA vocab, encoding, dataloaders |
train_final.py | Final training/eval pipeline |
model.py / train.py | Earlier v1 architecture (kept for reference) |
train_v2.py | Older grid-search script (outdated, lower metrics) |
If you use this model, cite the model:
@article{peptedgev2,
title={PeptEdgeV2: A Parameter-Efficient Multi-Scale Hybrid Architecture for Antimicrobial Peptide Classification},
author={Devansh Raulo},
year={2026}
}
If you use the dataset, cite its original authors:
@article{Pham2026GenPeptCurated2025,
title={GenPept-Curated-2025: A Benchmark Dataset for Antimicrobial Peptide Prediction with Homology-Controlled Partitioning},
author={Huynh Trong Pham and Bao Huynh and Thanh-Hoang Nguyen-Vo},
journal={bioRxiv},
year={2026},
doi={10.64898/2026.04.25.720793}
}
MIT