Downloads · 30 days
0
gschottlender/graphmatch-bioactivity-models
graphmatch-bioactivity-models is a machine learning model from gschottlender. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository is the versioned artifact store for graphmatch-bioactivity. It contains pretrained GraphMatch checkpoints and the frozen evaluation sets needed to reproduce the reported leave-one-protein-out (LOPO) ex…
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
58.8 MB
Likes
0
Public
Click a slice to open those files.
.gz42.9 MB · 73%
From the Hugging Face model README
This repository is the versioned artifact store for
graphmatch-bioactivity.
It contains pretrained GraphMatch checkpoints and the frozen evaluation sets
needed to reproduce the reported leave-one-protein-out (LOPO) experiments.
graphmatch-bioactivity-models-v1.tar.gz: versioned model registry.evaluation/v1/manifest.json: families, proteins, hashes, class counts and
compatible model identifiers.evaluation/v1/PFxxxxx/UNIPROT/test_balanced.parquet: 1:1 evaluation.evaluation/v1/PFxxxxx/UNIPROT/test_imbalanced_1to10.parquet: 1:10
evaluation for PR-AUC and early enrichment.The v1 evaluation release contains 25 held-out proteins from five Pfam families. Evaluation files are copied from the frozen experimental campaign without resampling. They must not be merged into training data when reporting LOPO performance.
pip install "graphmatch-bioactivity[models]"
graphmatch-bioactivity models download
graphmatch-bioactivity datasets list
graphmatch-bioactivity datasets download --family PF00069 --protein P28482
Use the model whose identifier has the same family and held-out protein as the test set. Thresholds stored in LOPO checkpoints were selected on validation, not on these test labels.
GraphMatch estimates shared bioactivity for molecular pairs; it does not receive a protein sequence during inference. Pseudo-negatives, bioactivity records and derived pairs inherit the biases and licensing conditions of their upstream sources. Attention and occlusion explain model behavior and are not proof of a pharmacophore or physical binding mode.
The source code is MIT-licensed. Model weights and evaluation data retain the terms and attribution requirements of their upstream data sources; consult the project documentation before redistribution or commercial use.