Downloads · 30 days
0
zxcyfr/Human-Protein-Language-Model-0.9
Human-Protein-Language-Model-0.9 is a machine learning model from zxcyfr. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as afl-3.0.
This repository contains 254 Pfam-level protein language model checkpoints generated for the study “Does domain-specific unsupervised fine-tuning improve protein language model performance?”.
Downloads · 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
662 GB
Likes
0
Public
Click a slice to open those files.
.pth662 GB · 100%
From the Hugging Face model README
This repository contains 254 Pfam-level protein language model checkpoints generated for the study “Does domain-specific unsupervised fine-tuning improve protein language model performance?”.
The checkpoints were derived from the ESM-2 650M model through unsupervised masked language modeling on sequences associated with human-proteome Pfam families. Protein sequences were clustered at a sequence-identity threshold of 0.9 before model fine-tuning.
Model checkpoints are organized by their corresponding Pfam identifiers. Each directory contains the checkpoint associated with that protein family or clan.
These checkpoints are provided for research on protein representation learning, protein language model benchmarking, and the evaluation of domain-specific unsupervised fine-tuning.
The models are intended for research use only and have not been validated for clinical or diagnostic applications.
The checkpoints can be loaded using the fair-esm package. Usage examples and benchmarking code are available in the project repository:
https://github.com/TianBoxue-lab/DS-UFT-Benchmark
The complete model collection is available at:
https://huggingface.co/collections/zxcyfr/ds-uft-protein-language-models-for-human-protein-families
This repository is distributed under the AFL-3.0 license.