Downloads · 30 days
40
18% of all-time downloads
virtual-human-chc/MolE
MolE is a machine learning model from virtual-human-chc. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
MolE learns task-independent molecular representations of chemicals via Graph Isomorphism Networks (GINs). Combined with an XGBoost classifier it estimates the probability of a compound inhibiting bacterial growth. Th…
Downloads · 30 days
40
18% of all-time downloads
All-time downloads
222
Public
Repo size
814 MB
Likes
0
Public
Click a slice to open those files.
.pth804 MB · 99%
From the Hugging Face model README
MolE learns task-independent molecular representations of chemicals via Graph Isomorphism Networks (GINs). Combined with an XGBoost classifier it estimates the probability of a compound inhibiting bacterial growth. The model was developed by Roberto Olayo Alarcon et al. and more information can be found in the GitHub repository and the accompanying paper.
MolE Antimicrobial Prediction (March 2024): Pretrained representation model + XGBoost classifier trained on antimicrobial screening data (Maier et al., 2018).
MolE is a non-contrastive self-supervised Graph Neural Network (GNN) framework that leverages unlabeled chemical structures to learn task-independent molecular representations. By combining a pre-trained MolE representation with experimentally validated compound-bacteria activity data, the project builds an antimicrobial prediction model that re-discovers recently reported growth-inhibitory compounds that are structurally distinct from current antibiotics. Using the model as a compound prioritization strategy, three human-targeted drugs are identified and experimentally confirm as growth-inhibitors of Staphylococcus aureus, highlighting MolE’s potential to accelerate the discovery of new antibiotics.
<!-- MolE integrates molecular graph-based representation learning with gradient-boosted decision trees for predicting antimicrobial potential. The approach involves: 1. **Representation learning:** A graph neural network (GINet) trained on 100,000 randomly sampled compounds to derive molecular embeddings from SMILES strings. 2. **Prediction:** These embeddings are used as input to an **XGBoost** model that predicts antimicrobial activity scores across 40 bacterial strains, based on data from *Maier et al., 2018*. The model was developed by **Roberto Olayo Alarcon et al.**. Further information is available in the [paper](https://www.nature.com/articles/s41467-025-58804-4). -->[n, 2], where n is the number of chemical compoundschem_name (str): Name of the molecule (e.g., Halicin).smiles (str): SMILES representation of the molecule (e.g.,C1=C(SC(=N1)SC2=NN=C(S2)N)[N+](=O)[O-]).input/examples_molecules.tsvmodel.pth, config.yaml).MolE-XGBoost-08.03.2024_14.20.pkl). The scores reflect the likelihood that a compound inhibits the growth of a given bacterial strain.[n × 40, 2], where n is the number of chemical compoundspred_id (str): Combination of the molecule name and bacterial strain (e.g., Halicin:Akkermansia muciniphila (NT5021)).antimicrobial_predictive_probability (float): Predicted probability that the compound inhibits microbial growth.output/example_molecules_prediction.tsvInstall the conda environment with all dependencies:
# Create the conda environment called virtual-human-chc-mole
conda env create -f environment.yaml
# Activate the environment
conda activate virtual-human-chc-mole
import torch
import yaml
import pickle
import pandas as pd
from huggingface_hub import hf_hub_download
from mole_package import ginet_concat, mole_antimicrobial_prediction, mole_representation, dataset_representation
class MolE:
def __init__(self, device='auto'):
repo = "virtual-human-chc/MolE"
self.device = "cuda:0" if device == "auto" and torch.cuda.is_available() else "cpu"
# Download + load
cfg = yaml.safe_load(open(hf_hub_download(repo, "config.yaml")))
self.model = ginet_concat.GINet(**cfg["model"]).to(self.device)
self.model.load_state_dict(torch.load(hf_hub_download(repo, "model.pth"), map_location=self.device))
self.xgb = pickle.load(open(hf_hub_download(repo, "MolE-XGBoost-08.03.2024_14.20.pkl"), "rb"))
def predict_from_smiles(self, smiles_tsv):
smiles_df = mole_representation.read_smiles(smiles_tsv, "smiles", "chem_name")
emb = dataset_representation.batch_representation(smiles_df, self.model, "smiles", "chem_name", device=self.device)
X_input = mole_antimicrobial_prediction.add_strains(
emb, "input/maier_screening_results.tsv.gz"
)
probs = self.xgb.predict_proba(X_input)[:, 1]
return pd.DataFrame(
{"antimicrobial_predictive_probability": probs},
index=X_input.index
)
# Run inference
mole = MolE()
pred = mole.predict_from_smiles("input/examples_molecules.tsv")
print(pred)
Code derived from https://github.com/rolayoalarcon/MolE is licensed under the MIT License, © 2024 Roberto Olayo Alarcon. Model weights are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0), © 2024 Roberto Olayo Alarcon. Additional code © 2025 Maksim Pavlov, licensed under MIT.