Downloads · 30 days
0
gggenwolai/uniamp
uniamp is a machine learning model from gggenwolai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
UniAMPv3 is a protein language model-based framework for antimicrobial peptide prediction and the identification of peptides with activity against Fusarium species.
Downloads · 30 days
0
Access
Public
Updated Aug 6, 2026
Repo size
14 GB
Likes
0
Public
Click a slice to open those files.
.bin11.3 GB · 81%
From the Hugging Face model README
UniAMPv3 is a protein language model-based framework for antimicrobial peptide prediction and the identification of peptides with activity against Fusarium species.
This repository contains the sequence datasets, predefined cross-validation partitions and pretrained backbone weights used for model development and evaluation. The final trained UniAMPv3 checkpoint will be added upon acceptance of the associated manuscript.
.
├── data/
│ ├── AMPs/
│ ├── Fusarium_AMPs/
│ └── Fusarium_AMPs_GraphPart/
├── checkpoint/
├── prott5_pre_models/
├── unirep_pre_models/
├── esmc_600m_2024_12_v0.pth
└── README.md
The checkpoint/ directory will be added when the final trained model is publicly released.
The data/ directory contains three predefined five-fold datasets.
data/AMPs/ contains the general antimicrobial peptide dataset used for model training and evaluation.
data/Fusarium_AMPs/ contains the anti-Fusarium peptide dataset used for pathogen-specific model training and evaluation.
data/Fusarium_AMPs_GraphPart/ contains a more stringent five-fold partition of the anti-Fusarium peptide dataset generated using GraphPart. Sequences assigned to different folds have less than 40% pairwise sequence identity.
Each dataset is divided into fold0–fold4, corresponding to the five predefined cross-validation folds used in the associated study.
Dataset files contain amino acid sequences, class labels and partition information where applicable. Only the 20 standard amino acids are retained in the processed sequence datasets.
The repository includes the original pretrained weights used to initialize the three protein language model backbones before domain-adaptive pretraining:
prott5_pre_models/: pretrained ProtT5 model files;unirep_pre_models/: pretrained UniRep model files;esmc_600m_2024_12_v0.pth: pretrained ESMC-600M model weights.These files are not UniAMPv3-trained checkpoints and have not undergone domain-adaptive pretraining in this repository. They are included to facilitate reproducible model initialization.
The pretrained backbone models remain subject to the licenses and terms specified by their original developers and distributors.
The final UniAMPv3 model is based on full-parameter domain-adaptive pretraining of ProtT5 followed by LoRA-based supervised fine-tuning for anti-Fusarium peptide prediction.
The trained checkpoint and its associated configuration files will be released in the checkpoint/ directory upon acceptance of the associated manuscript.
The source code used for data processing, model training, evaluation and inference will be publicly released on GitHub upon acceptance of the associated manuscript.
GitHub repository: to be added.
The sequence datasets and predefined data partitions used for model development and evaluation are publicly available in this repository.
The processed datasets and trained UniAMPv3 checkpoint are provided under the licenses specified in this repository. Data and pretrained model files obtained from external sources remain subject to the licenses and terms of their original providers.
Questions regarding the datasets or model should be directed to the corresponding author of the associated manuscript.