Downloads · 30 days
0
ibm-research/otter_dude_classifier
otter_dude_classifier is a machine learning model from ibm-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Otter models are based on Graph Neural Networks (GNN) that propagates initial embeddings through a set of layers that upgrade input embedding according to the node neighbours. The architecture of GNN consists of two m…
Downloads · 30 days
0
Access
Public
Updated Jun 26, 2023
Repo size
3.6 MB
Likes
3
Public
Click a slice to open those files.
.pt1.8 MB · 100%
From the Hugging Face model README
Otter models are based on Graph Neural Networks (GNN) that propagates initial embeddings through a set of layers that upgrade input embedding according to the node neighbours. The architecture of GNN consists of two main blocks: encoder and decoder.
Model type:
For link prediction, we consider three choices of scoring functions: DistMult, TransE and a Binary Classifier that are commonly used in the literature. The outcomes of scoring of each triple are then compared against actual labels using negative log likelihood loss function.
| Scoring Type | Noisy Links | Flow Control | Regression |
|---|---|---|---|
| Classifier Head | No | Yes | No |
Model training data:
The model was trained over a preprocessed version of DUDe. Our preprocessed version of DUDe includes 1,452,568 instances of drug-target interactions. To prevent any data leakage, we eliminated the negative interactions and the overlapping triples with the TDC DTI dataset. As a result, we were left with a total of 40,216 drug-target interaction pairs.
Model results:
<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} .tg .tg-0pky{border-color:inherit;text-align:center;vertical-align:centr;text-emphasis:bold} </style> <table class="tg"> <thead> <tr> <th class="tg-0pky">Dataset</th> <th class="tg-c3ow">DTI DG</th> <th class="tg-c3ow" colspan="3">DAVIS</th> <th class="tg-c3ow" colspan="3">KIBA</th> </tr> </thead> <tbody> <tr> <td class="tg-0pky">Splits</td> <td class="tg-c3ow">Temporal</td> <td class="tg-c3ow">Random</td> <td class="tg-c3ow">Target</td> <td class="tg-c3ow">Drug</td> <td class="tg-c3ow">Random</td> <td class="tg-c3ow">Target</td> <td class="tg-c3ow">Drug</td> </tr> <tr> <td class="tg-0pky">Results</td> <td class="tg-c3ow">0.579</td> <td class="tg-c3ow">0.808</td> <td class="tg-c3ow">0.574</td> <td class="tg-c3ow">0.167</td> <td class="tg-c3ow">0.860</td> <td class="tg-c3ow">0.641</td> <td class="tg-c3ow">0.630</td> </tr> </tbody> </table>Paper or resources for more information:
License:
MIT
Where to send questions or comments about the model:
Clone the repo:
git clone https://github.com/IBM/otter-knowledge.git
cd otter-knowledge
Run the inference for Proteins:
Replace test_data with the path to a CSV file containing the protein sequences, name_of_the_column with the name of the column of the protein sequence in the CSV and output_path with the filename of the JSON file to be created with the embeddings.
python inference.py --input_path test_data --sequence_column name_of_the_column --model_path ibm/otter_dude_classifier --output_path output_path
Run the inference for Drugs:
Replace test_data with the path to a CSV file containing the Drug SMILES, name_of_the_column with the name of the column of the SMILES in the CSV and output_path with the filename of the JSON file to be created with the embeddings..*
python inference.py --input_path test_data --sequence_column name_of_the_column input_type Drug --relation_name smiles --model_path ibm/otter_dude_classifier --output_path output_path