Downloads · 30 days
0
nvidia/NV-KERMT-70M-v2
NV-KERMT-70M-v2 is a graph machine learning model from nvidia. Use it for the graph machine learning task on the model card, and read the license before you ship it in a product. It is set up for bionemo. The card lists the license as other.
Source code, training scripts, and inference utilities for this model: github.com/NVIDIA-BioNeMo/KERMT (v2.0 branch / v2.0.0 release tag)
Downloads · 30 days
0
Access
Public
Updated Jun 23, 2026
Repo size
282 MB
Likes
7
Public
Click a slice to open those files.
.pt282 MB · 100%
From the Hugging Face model README
Source code, training scripts, and inference utilities for this model: github.com/NVIDIA-BioNeMo/KERMT (v2.0 branch / v2.0.0 release tag)
Contrastive KERMT (Kinetic GROVER Multi-Task) is a graph-transformer foundation model pretrained to learn chemically meaningful molecular representations for downstream ADMET (absorption, distribution, metabolism, excretion, toxicity) property prediction in drug discovery. The model encodes a 2D molecular graph into a latent representation under a single joint probabilistic objective that combines SMILES reconstruction, in-batch contrastive discrimination, and chemistry-specific self-supervision (atom-context, bond-context, and functional group prediction), all formulated as unit-weighted log-probability factors. The released checkpoint was pretrained for 100 epochs on a corpus combining an 11M-molecule ZINC15+ChEMBL base pool (following the pretraining-data protocol of Rong et al. 2020) with Biogen ADMET, ExpansionRX, and ChEMBL-MT (~125K additional molecules), and is intended as a starting point for downstream multi-task ADMET fine-tuning. Contrastive KERMT was developed by NVIDIA as part of the KERMT v2.0 release. This model is ready for commercial or non-commercial use.<br>
Copyright © 2026, NVIDIA Corporation. All rights reserved.<br>
The source code is made available under Apache License, Version 2.0. See LICENSE in the source repository at https://github.com/NVIDIA-BioNeMo/KERMT.<br>
The model weights are made available under the NVIDIA Open Model License.<br>
Global<br>
Computational chemistry and machine-learning researchers in drug discovery — particularly those working on ADMET / Drug Metabolism and Pharmacokinetics (DMPK) prediction — who need a pretrained molecular graph encoder that can be fine-tuned on multi-endpoint ADMET datasets, used as a feature extractor for property-prediction pipelines, or studied as a baseline in molecular-representation-learning research. The released checkpoint is a pretrained backbone; users are expected to fine-tune it on their own labeled datasets for specific ADMET endpoints before using predictions in downstream workflows.<br>
NGC 06/10/2026 via https://catalog.ngc.nvidia.com/orgs/nvidia/teams/clara/resources/kermt-contrastive <br> Hugging Face 06/10/2026 via https://huggingface.co/nvidia/NV-KERMT-70M-v2 <br>
Architecture Type: Transformer (graph-transformer with local message passing + global self-attention) <br>
Network Architecture: KERMT graph-transformer encoder (extension of GROVER) with a probabilistic latent head, an in-batch contrastive auxiliary variable, a SMILES-reconstruction transformer decoder, and chemistry-specific vocabulary prediction heads. Encoder: hidden size 800, 6 message-passing-plus-attention layers, 4 attention heads per layer, 1 multi-task (MT) block, PReLU activation, dropout 0.1. Decoder: 3 transformer layers, 8 attention heads, 512 hidden / latent dimension, FFN hidden 2048, rotary positional encoding (RoPE). Latent dimension 512. <br>
This model was developed based on KERMT (Adrian et al. 2025, arXiv:2510.12719), in turn based on GROVER (Rong et al. 2020). <br>
Number of model parameters: 7.06 × 10^7 <br>
Input Type(s): Text (SMILES string representing a 2D molecular structure) <br>
Input Format(s): UTF-8 SMILES (Simplified Molecular Input Line Entry System) <br>
Input Parameters: One-Dimensional (1D) text <br>
Other Properties Related to Input: The input is a canonical SMILES string parseable by RDKit (an open-source cheminformatics toolkit); molecules are internally featurized into 2D atom-and-bond graphs prior to encoding. Recommended maximum sequence length for the SMILES decoder is 512 tokens (the value used at pretraining time); molecules whose canonical SMILES exceed this length should be truncated or omitted. Inputs are not text in the natural-language sense and are not subject to natural-language preprocessing (no tokenization in the human-language sense; characters are mapped via a chemistry-specific tokenizer matching the bundled SMILES vocabulary). <br>
Output Type(s): Numerical tensors (molecular embeddings) and, when downstream task-specific heads are present, scalar ADMET property predictions. Optionally, generated SMILES strings via the pretraining-time SMILES decoder. <br>
Output Format(s): <br>
Output Parameters: One-Dimensional (1D) embedding / prediction vectors. <br>
Other Properties Related to Output: Embeddings are intended as inputs to downstream property-prediction heads, similarity computations, or visualization (PCA / t-SNE / UMAP). Predictions are statistical estimates derived from training data and should not be used as substitutes for experimental measurement in safety-critical drug development decisions.
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions. <br>
Runtime Engine(s):
Supported Hardware Microarchitecture Compatibility: <br> NVIDIA GPU with compute capability 7.0 (Volta) or newer is recommended; at least 32 GB GPU vRAM is recommended for pretraining and fine-tuning workloads (inference uses less). The following microarchitecture families are supported: <br>
Supported Operating System(s):
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment. <br>
kermt-contrastive v2.0 — Pretrained Contrastive KERMT checkpoint trained for 100 epochs on the pooled corpus described under "Training Dataset" below (11M ZINC15+ChEMBL base pool + Biogen ADMET + ExpansionRX + ChEMBL-MT). Inference-only release (training-time optimizer state stripped). <br>
The full KERMT v2.0 release includes (a) source code in the public repository at https://github.com/NVIDIA-BioNeMo/KERMT (Apache 2.0), (b) the pretrained checkpoint described here, and (c) the bundled pretraining vocabulary files (atom vocab JSON, bond vocab JSON, SMILES vocab pickle) required to reproduce the model's tokenization. <br>
Not Applicable
Less than 100 Million Datapoints. Pretraining corpus contains approximately 11.1 million unique canonical SMILES strings after pooling and RDKit-based deduplication: an 11M ZINC15+ChEMBL base pool (per Rong et al. 2020) plus Biogen ADMET (~3.5K), ExpansionRX (~7.6K), and ChEMBL-MT (~114K). <br>
** Non-Audio, Image, Text Training Data Size: ~11.1 × 10^6 molecules (SMILES strings); total on-disk text size ≈ a few hundred MB depending on serialization. <br>
** Data Collection Method by dataset: Hybrid: Automated, Manually-Collected
** Labeling Method by dataset: Not Applicable — pretraining is unsupervised and uses no human-provided labels. Targets (atom-context, bond-context, functional groups) are computed on the fly from canonical SMILES using deterministic RDKit-based rules; SMILES-reconstruction targets are the input SMILES themselves.
Properties: ~11M molecular SMILES strings; the data are text representations of 2D chemical structures and contain no personal data, copyrighted natural-language content, or human-language linguistic content. All molecules are deduplicated by canonical SMILES. The training pool is scaffold-balanced via Bemis-Murcko scaffolds for the train/val split. <br>
Pretraining is fully unsupervised and uses no dedicated testing dataset (no held-out test split at the pretraining stage). Model quality is assessed downstream via the fine-tuning Evaluation Datasets below.
Data Collection Method: Not Applicable <br> Labeling Method: Not Applicable <br> Properties: Not Applicable.
Benchmark Score:
Downstream evaluation is performed by fine-tuning the released pretrained checkpoint on three independent ADMET benchmarks and reporting Mean Absolute Error (MAE), Pearson correlation coefficient (r), and Spearman correlation coefficient (ρ) per endpoint, averaged over multiple random seeds:
Data Collection Method by dataset: Hybrid: Manually-Collected, Automated
Labeling Method by dataset:
Properties: Continuous-valued ADMET endpoint measurements (intrinsic clearance, permeability, solubility, plasma protein binding, hERG inhibition, etc.) from in vitro and in vivo assays. Data are scalar regression targets per molecule; no images, video, or natural-language content. ChEMBL-MT contributes the toxicity endpoint (hERG inhibition); the ADME endpoints come from all three benchmarked datasets.
Acceleration Engine: PyTorch (the released checkpoint is loadable via the KERMT codebase). <br> Test Hardware: <br>
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
For more detailed information on ethical considerations for this model, please see the Model Card++ Bias, Explainability, Safety & Security, and Privacy Subcards. <br>
Users are responsible for ensuring the physical properties of model-generated molecules are appropriately evaluated and comply with applicable safety regulations and ethical standards. <br>
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here. <br>