Downloads · 30 days
0
catly/scKGBERT
scKGBERT is a machine learning model from catly. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated May 7, 2025
Repo size
70.7 MB
Likes
0
Public
Click a slice to open those files.
.bin70.7 MB · 100%
From the Hugging Face model README
This repository is the official implementation of scKGBERT.
<hr>scKGBERT is the first single-cell foundation model to integrate genetic knowledge graph pre-training and RNA sequence pre-training into a unified framework. The model utilizes a Gaussian attention mechanism within the transformer architecture to align key knowledge with RNA sequence representations, enhancing its interpretability in learning gene and cell profiles. This approach allows scKGBERT to understand and learn both the biological semantics and genetic contexts of genes, capturing extensive interactions between independent genes and dynamic network information within cells.
<div align=center><img src="./fig/overall.png" style="zoom:50%;" /> </div>To run our code, please install dependency packages.
huggingface_hub 0.15.1
networkx 2.8.4
numpy 1.24.3
python 3.8.17
scikit-learn 1.2.2
tokenizers 0.13.2
torch 1.13.0+cu117
transformers 4.31.0
This project mainly contains the following parts.
├── models # pre-trained scKGBERT model.
├── utils # Utility classes folder
│ ├── KG_TransR.py # TransR training code.
│ ├── pretrainer.py # The Pretrainer code is used to train a model on large-scale datasets to obtain general feature representations.
│ ├── scKGBERT_model.py # scKGBERT model.
│ ├── tokenizer.py # The Tokenizer class is used to segment data into units and convert them into numerical forms that a model can process.
│ └── utils_pkg.py # Other utility classes used.
├── pretrainscKGBERT.py # The scKGBERT pre-training code.
If you want to use our pre-trained model directly for molecular property prediction, please run the following command:
>> deepspeed --num_gpus=2 pretrainscKGBERT.py
The processed single-cell transcriptome data can be used with our model to accomplish the following tasks:
Gene annotation
Cell and cell subtype annotation
Cancer drug response
biomarker analysis