Downloads · 30 days
12
67% of all-time downloads
OneScience-Group/La-Proteina
La-Proteina is a machine learning model from OneScience-Group. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="center" <strong <span style="font-size: 30px;"La-Proteina</span </strong </p
Downloads · 30 days
12
67% of all-time downloads
All-time downloads
18
Public
Repo size
108 KB
Likes
0
Public
Click a slice to open those files.
.pyc710 KB · 52%
From the Hugging Face model README
La-Proteina is a protein structure generation model based on Partially Latent Flow Matching. It can directly generate all-atom protein structures together with their corresponding amino acid sequences.
Paper: La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching (arXiv 2025).
La-Proteina explicitly models the protein backbone (backbone CA), while sequence and atom-level details are captured through fixed-dimensional latent variables for each residue. This effectively avoids the challenges introduced by explicit side-chain representations.
| Use case | Description |
|---|---|
| Protein structure generation | Generates protein backbones (backbone CA) and local latents based on flow matching. |
| Motif-constrained generation | Supports motif position and sequence constraints for backbone design around functional motifs. |
| Diffusion model training | Trains the main La-Proteina model on PDB datasets. |
| Autoencoder training and inference | Trains the local-latent autoencoder and performs encoding, decoding, and reconstruction evaluation on PDB structures. |
| Generated result evaluation | Computes metrics such as RMSD, sequence recovery, and (co-)designability. |
You can try intelligent one-click AI4S programming through the OneCode online environment:
Try intelligent one-click AI4S programming
Hardware Requirements
Software Requirements
DCU users who want to learn more about adaptation details can contact [email protected].
Environment Checks
nvidia-smi
hy-smi
Optional: override the default paths through environment variables:
export LAPROTEINA_ROOT=/path/to/la-proteina
export LAPROTEINA_DATASET_DIR=/path/to/dataset
export LAPROTEINA_CHECKPOINTS_DIR=/path/to/checkpoints_laproteina
export DATA_PATH=/path/to/dataset
conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[bio] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
If the following code cannot find required libraries at runtime, activate CUDA as shown below.
source ${ROCM_PATH}/cuda/env.sh
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib/python3.11/site-packages/fastpt/torch/lib:$LD_LIBRARY_PATH"
# By default, the package is downloaded to the model folder under the current path. To change this, adjust the path after local_dir.
hf download OneScience-Group/La-Proteina --local-dir ./model
cd model
They will be uploaded to Hugging Face soon, and command-line downloads will be supported later.
It is recommended to run the scripts under the laproteina directory so that outputs are managed in one place.
run_train.sh)bash scripts/run_train.sh
Common Hydra parameter overrides:
# Single-device debugging
bash scripts/run_train.sh hardware.ngpus_per_node_=1 single=true
# Specify a run name
bash scripts/run_train.sh run_name=my_laproteina_run
# Override dataset or network configuration (for motif training scenarios)
bash scripts/run_train.sh dataset=pdb/pdb_train_motif_aa nn=local_latents_score_nn_160M_motif_idx_aa
Output:
./store/<run_name>/: training logs, checkpoints, and Hydra configurationsrun_generate.sh)bash scripts/run_generate.sh
By default, the inference_ucond_tri configuration is used for unconditional generation (LD2 model + AE1).
Switching generation configurations:
# Unconditional generation without triangle attention
bash scripts/run_generate.sh --config_name inference_ucond_notri
# Unconditional generation for long chains (300-800 residues)
bash scripts/run_generate.sh --config_name inference_ucond_notri_long
# Indexed all-atom motif scaffolding
bash scripts/run_generate.sh --config_name inference_motif_idx_aa
# Indexed tip-atom motif scaffolding
bash scripts/run_generate.sh --config_name inference_motif_idx_tip
# Unindexed all-atom motif scaffolding
bash scripts/run_generate.sh --config_name inference_motif_uidx_aa
# Unindexed tip-atom motif scaffolding
bash scripts/run_generate.sh --config_name inference_motif_uidx_tip
Output:
./inference/<config_name>/: generated protein structure files and metadatarun_evaluate.sh)bash scripts/run_evaluate.sh
By default, this evaluates the generated results corresponding to the inference_ucond_tri configuration.
Preparing ProteinMPNN weights:
ProteinMPNN weights must be downloaded before evaluation. They can be downloaded from the Hugging Face community:
hf download OneScience-Group/ProteinMPNN --local-dir ./weight
Output:
./inference/<config_name>/evaluation/: evaluation result filesrun_ae_infer.sh)bash scripts/run_ae_infer.sh
This performs encode-decode reconstruction on the PDB dataset and evaluates reconstruction metrics, such as all-atom RMSD and sequence recovery.
The script checks whether DATA_PATH/pdb_train and AE1_ucond_512.ckpt exist.
The script internally calls:
python infer_laproteina_ae.py "$@"
Common parameter overrides:
| Environment variable | Default value | Description |
|---|---|---|
LAPROTEINA_ROOT | ${ONESCIENCE_DATASETS_DIR}/la-proteina | Root directory for data and weights |
LAPROTEINA_CHECKPOINTS_DIR | ${LAPROTEINA_ROOT}/checkpoints_laproteina | Autoencoder weight directory |
DATA_PATH | ${LAPROTEINA_ROOT}/dataset | Dataset directory |
Output:
./inference_ae/: reconstructed structures and evaluation metricsONESCIENCE_DATASETS_DIR environment variable is correctly set before running the scripts.DATA_PATH/pdb_train and AE1_ucond_512.ckpt exist by default. If either is missing, it will report an error and exit.dataset=pdb is available. dataset=genie2 and dataset=pdb_multimer are not packaged in the current repository snapshot, and running them will produce an explicit error.LD_LIBRARY_PATH values and can run directly on Hygon DCU platforms. On CUDA platforms, these settings can be ignored or adjusted as needed.script_utils/download_pmpnn_weights.sh in advance to download them.examples/biosciences/laproteina so that output directories remain unified.+CK_PATH=... to specify the checkpoint root path. If it is not provided, the script automatically sets it to LAPROTEINA_ROOT.| Platform | OneScience main repository | Skills repository |
|---|---|---|
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
The example code is licensed under Apache 2.0. La-Proteina model weights are licensed under the NVIDIA Open Model License Agreement, and other materials are licensed under CC-BY 4.0.
If you use La-Proteina in your research, please cite the original paper:
@article{geffner2025laproteina,
title={La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching},
author={Geffner, Tomas and Didi, Kieran and Cao, Zhonglin and Reidenbach, Danny and Zhang, Zuobai and Dallago, Christian and Kucukbenli, Emine and Kreis, Karsten and Vahdat, Arash},
journal={arXiv preprint arXiv:2507.09466},
year={2025}
}