Downloads · 30 days
0
ibm-research/otter_stitch_classifier
otter_stitch_classifier is a machine learning model from ibm-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Otter models are based on Graph Neural Networks (GNN) that propagates initial embeddings through a set of layers that upgrade input embedding according to the node neighbours. The architecture of GNN consists of two m…
Downloads · 30 days
0
Access
Public
Updated Jun 26, 2023
Repo size
1.8 MB
Likes
2
Public
Click a slice to open those files.
.pt1.8 MB · 100%
From the Hugging Face model README
Otter models are based on Graph Neural Networks (GNN) that propagates initial embeddings through a set of layers that upgrade input embedding according to the node neighbours. The architecture of GNN consists of two main blocks: encoder and decoder.
Model type:
For link prediction, we consider three choices of scoring functions: DistMult, TransE and a Binary Classifier that are commonly used in the literature. The outcomes of scoring of each triple are then compared against actual labels using negative log likelihood loss function.
| Scoring Type | Noisy Links | Flow Control | Regression |
|---|---|---|---|
| Classifier Head | No | Yes | No |
Model training data:
The model was trained over STITCH. STITCH (Search Tool for Interacting Chemicals) is a database of known and predicted interactions between chemicals represented by SMILES strings and proteins whose sequences are taken from STRING database. It contains 10,717,791 triples for 17,572 different chemicals and 1,886,496 different proteins. Furthermore, the graph was split into 5 roughly same size subgraphs and GNN was trained sequentially on each of them by upgrading the model trained using the previous subgraph.
Model results:
<style type="text/css"> .tg {border-collapse:collapse;border-spacing:0;} .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; overflow:hidden;padding:10px 5px;word-break:normal;} .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px; font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;} .tg .tg-c3ow{border-color:inherit;text-align:center;vertical-align:top} .tg .tg-0pky{border-color:inherit;text-align:center;vertical-align:centr;text-emphasis:bold} </style> <table class="tg"> <thead> <tr> <th class="tg-0pky">Dataset</th> <th class="tg-c3ow">DTI DG</th> <th class="tg-c3ow" colspan="3">DAVIS</th> <th class="tg-c3ow" colspan="3">KIBA</th> </tr> </thead> <tbody> <tr> <td class="tg-0pky">Splits</td> <td class="tg-c3ow">Temporal</td> <td class="tg-c3ow">Random</td> <td class="tg-c3ow">Target</td> <td class="tg-c3ow">Drug</td> <td class="tg-c3ow">Random</td> <td class="tg-c3ow">Target</td> <td class="tg-c3ow">Drug</td> </tr> <tr> <td class="tg-0pky">Results</td> <td class="tg-c3ow">0.576</td> <td class="tg-c3ow">0.804</td> <td class="tg-c3ow">0.571</td> <td class="tg-c3ow">0.156</td> <td class="tg-c3ow">0.856</td> <td class="tg-c3ow">0.627</td> <td class="tg-c3ow">0.585</td> </tr> </tbody> </table>Paper or resources for more information:
License:
MIT
Where to send questions or comments about the model:
Clone the repo:
git clone https://github.com/IBM/otter-knowledge.git
cd otter-knowledge
Run the inference for Proteins:
Replace test_data with the path to a CSV file containing the protein sequences, name_of_the_column with the name of the column of the protein sequence in the CSV and output_path with the filename of the JSON file to be created with the embeddings.
python inference.py --input_path test_data --sequence_column name_of_the_column --model_path ibm/otter_stitch_classifier --output_path output_path
Run the inference for Drugs:
Replace test_data with the path to a CSV file containing the Drug SMILES, name_of_the_column with the name of the column of the SMILES in the CSV and output_path with the filename of the JSON file to be created with the embeddings..*
python inference.py --input_path test_data --sequence_column name_of_the_column input_type Drug --relation_name smiles --model_path ibm/otter_stitch_classifier --output_path output_path