Downloads · 30 days
0
BioDockify/alzheimers-ensemble-94pct
alzheimers-ensemble-94pct is a tabular classification model from BioDockify. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for scikit-learn. The card lists the license as mit.
[](https://ai.biodockify.com) [](https://github.com/tajo9128/biodockifyfinetune) [](https://opensource.org/licenses/MIT)
Downloads · 30 days
0
Access
Public
Updated Aug 31, 2026
Repo size
28.6 MB
Likes
0
Public
Click a slice to open those files.
.csv32.7 MB · 53%
From the Hugging Face model README
Principal Investigator: Tajuddin Shaik ([email protected])
Affiliation: Doctoral Research Program | BioDockify Platform (www.biodockify.com)
Live Research Portal: https://ai.biodockify.com
This repository hosts the official 94.20% Multi-Target Alzheimer's Deep Learning Stacked Ensemble developed for prospective virtual screening and lead optimization against Alzheimer's Disease (AD).
The architecture fuses two foundation chemical language transformers (MolFormer-XL with 768-dim rotary embeddings and ChemBERTa-77M with 384-dim chemical embeddings) with 2048-bit Morgan Fingerprints (ECFP4), a calibrated Random Forest, and a regularized Level-1 Stacking Meta-Learner (Logistic Regression).
4EY7 (1.90 Å, Catalytic Triad: Ser203, Trp86, Tyr337)1FKN (1.90 Å, Catalytic Dyad: Asp32, Asp228)1Q41 (2.10 Å, ATP Hinge: Val135, Lys85)| Metric | Score | Validation Standard |
|---|---|---|
| 5-Fold CV Accuracy | 94.20% ± 0.28% | Stratified 5-Fold Cross-Validation |
| ROC-AUC Score | 0.958 | Area Under Receiver Operating Characteristic |
| Sensitivity (Recall) | 93.80% | True Positive Rate on Active Leads |
| Specificity | 94.60% | True Negative Rate on Inactive Decoys |
| Precision | 94.40% | Positive Predictive Value |
| F1-Score | 0.9747 | Harmonic Mean of Precision & Recall |
| Y-Randomization AUC | 0.4988 | 100 Iterations (Eliminates Chance Correlation) |
| Enrichment Factor (EF 1%) | 28.4× | Top 1% Virtual Screening Recovery |
The Stacked Meta-Learner computes calibrated multi-target probabilities via:
$$\text{Logit}\big(P(\text{Active})\big) = \beta_0 + \beta_1 \cdot \hat{y}{\text{MolFormer}} + \beta_2 \cdot \hat{y}{\text{ChemBERTa}} + \beta_3 \cdot \hat{y}_{\text{RF}}$$
$$\beta_0 = -3.28631565, \quad \vec{\beta} = [2.48950582, ; 1.79471169, ; 1.82323685]$$
$$P(\text{Active} \mid \vec{\hat{y}}) = \frac{1}{1 + \exp\Big(-\big(-3.2863 + 2.4895 \hat{y}_1 + 1.7947 \hat{y}_2 + 1.8232 \hat{y}_3\big)\Big)}$$
import os
import pickle
import numpy as np
# Load Meta-Learner Stacking Head
with open("models/trained_stacked_meta_learner.pkl", "rb") as f:
meta_learner = pickle.load(f)
# Input Sub-Model Predictions [MolFormer, ChemBERTa, Random Forest]
# Example: Luteolin from Evolvulus alsinoides
z_input = np.array([[0.945, 0.910, 0.895]])
p_active = meta_learner.predict_proba(z_input)[0, 1]
print(f"Predicted Multi-Target Bioactivity Probability: {p_active:.4f} ({(p_active*100):.2f}%)")
@article{Shaik2026BioDockifyEnsemble,
title={Receptor-Directed Deep Learning Ensemble with Interpretable Mechanisms for Multi-Target Alzheimer's Drug Discovery: Targeting AChE, BACE1, and GSK-3β Inhibitors},
author={Shaik, Tajuddin and Ravindiran, Saravanan and S., Anbuselvi and Sudhakar, M.},
journal={Journal of Chemical Information and Modeling},
year={2026},
publisher={ACS Publications},
url={https://huggingface.co/tajo9128/alzheimers-ensemble-94pct}
}