Downloads · 30 days
23
5% of all-time downloads
genbio-ai/GB.StructureTokenizer
GB.StructureTokenizer is a machine learning model from genbio-ai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
[](https://github.com/genbio-ai/ModelGenerator/blob/main/LICENSE)
Downloads · 30 days
23
5% of all-time downloads
All-time downloads
468
Public
Parameters
275M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
From the Hugging Face model README
GB.StructureTokenizer is a VQ-VAE-based tokenizer designed for protein structure prediction and tokenization. It encodes amino-acid-agnostic backbone structures into discrete tokens and reconstructs the full atomic-level structures, including side chains. This tokenizer facilitates the integration of 3D protein structure data with sequence-based language models, enabling efficient and accurate multimodal protein modeling.

GB.StructureTokenizer is built on a Vector Quantized Variational Autoencoder (VQ-VAE) architecture with the following components:
This model strikes a balance between reconstruction fidelity and structural locality, optimizing its suitability for downstream tasks such as structure prediction, homology detection, and multimodal protein modeling.



Please see experiments/GB.StructureTokenizer in Model Generator for more details.
Install Model Generator
To reproduce the reconstruction results in the paper, we provide a preprocessed CASP15 dataset at genbio-ai/sample-structure-dataset. It could be downloaded via
huggingface-cli download genbio-ai/sample-structure-dataset --repo-type dataset --local-dir ./data/protstruct_sample_data/
This dataset is based on the CASP15 dataset, which can be referenced at:
The downloaded directory includes:
registries folder containing a CSV file with metadata such as filenames and PDB IDs.CASP15_merged folder containing PDB files, where domains are merged in the same way as described in Bhattacharya-Lab/CASP15.To use customized data, you can prepare a dataset with the following structure:
cif.gz, cif, ent.gz, pdb).Then, you need to prepare a registry file in CSV format using the following command:
python experiments/GB.StructureTokenizer/register_dataset.py \
--folder_path /path/to/folder_path \
--format cif.gz \
--output_file /path/to/output_file.csv
You need to replace the folder_path and the registry_path in the following steps accordingly.
If you use the provided CASP15 dataset, you can run the combined encoding and decoding task using the following command:
CUDA_VISIBLE_DEVICES=0 mgen predict --config=experiments/GB.StructureTokenizer/encode_decode.yaml
If you use your own dataset, you need to update the folder_path and the registry_path in the encode_decode.yaml configuration file or override them when running the command. Example:
CUDA_VISIBLE_DEVICES=0 mgen predict --config experiments/GB.StructureTokenizer/encode_decode.yaml \
--data.init_args.config.proteins_datasets_configs.name="your_dataset_name" \
--data.init_args.config.proteins_datasets_configs.registry_path="your_dataset_folder_path" \
--data.init_args.config.proteins_datasets_configs.folder_path="your_dataset_registry_path" \
--trainer.callbacks.dict_kwargs.output_dir="your_output_dir"
The input and the output can be summarized as follows:
Input:
Output:
logs/protstruct_model/.output.pdb.input.pdb.Notes:
We use VS Code + Protein Viewer Extension to visualize the protein structures. It's a beginner-friendly tool for VS Code users. You could also use your preferred protein structure viewer to visualize the structures (e.g., PyMOL, ChimeraX, etc.), but here we focus on this extension.
If you have run the Running Encoding and Decoding Task, you could find the decoded structures and their corresponding original structures in the output directory. You could visualize them as follows.
output.pdb and input.pdb pair in the side panel. Select both files when holding the Ctrl key (for Mac users, hold the Cmd key). 


Please cite GB.StructureTokenizer using the following BibTex code:
@inproceedings{zhang_balancing_2024,
title = {Balancing Locality and Reconstruction in Protein Structure Tokenizer},
url = {https://www.biorxiv.org/content/10.1101/2024.12.02.626366v2},
doi = {10.1101/2024.12.02.626366},
publisher = {bioRxiv},
author = {Zhang, Jiayou and Meynard-Piganeau, Barthelemy and Gong, James and Cheng, Xingyi and Luo, Yingtao and Ly, Hugo and Song, Le and Xing, Eric},
year = {2024},
booktitle={NeurIPS 2024 Workshop on Machine Learning in Structural Biology (MLSB)},
}