Downloads · 30 days
106
8% of all-time downloads
heqin-zhu/structRFM
structRFM is a machine learning model from heqin-zhu. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
106
8% of all-time downloads
All-time downloads
1.4K
Public
Parameters
86.1M
344 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors344 MB · 100%
From the Hugging Face model README
structRFM is a fully open-source structure-guided RNA foundation model that integrates sequence and structural knowledge through innovative pre-training strategies. By leveraging 21 million sequence-structure pairs and a novel Structure-guided Masked Language Modeling (SgMLM) approach, structRFM achieves state-of-the-art performance across a broad spectrum of RNA structural and functional inference tasks, setting new benchmarks for reliability and generalizability.
<div align="center"> <img src="images/Fig1.png", width="800"> <sub>Figure: Overview of architecture and downstream applications</sub> </div>Install package: pip install transformers
import os
from transformers import AutoModel, AutoTokenizer
model_path = 'heqin-zhu/structRFM'
# model_path = os.getenv('structRFM_checkpoint')
model = AutoModel.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path)
# single sequence
seq = 'GUCCCAACUCUUGCGGGGAGGGAU'
inputs = tokenizer(seq, return_tensors="pt")
outputs = model(**inputs)
print('>>> single seq, length:', len(seq))
for k, v in outputs.items():
print(k, v.shape)
print(outputs.last_hidden_state.shape)
# batch mode
seqs = ["GUCCCAA", 'AGUGUUG', 'AUGUAGUTCUN']
inputs = tokenizer(
seqs,
add_special_tokens=True,
max_length=512,
padding='max_length',
truncation=True,
return_tensors='pt'
)
outputs = model(**inputs) # note that the output sequential features are padded to max-length
print('>>> batch seqs, batch:', len(seqs))
for k, v in outputs.items():
print(k, v.shape)
'''
>>> single seq, length: 24
last_hidden_state torch.Size([1, 24, 768])
pooler_output torch.Size([1, 768])
torch.Size([1, 24, 768])
>>> batch seqs, batch: 3
last_hidden_state torch.Size([3, 512, 768])
pooler_output torch.Size([3, 768])
'''
pip install transformers structRFM BPfold
wget https://github.com/heqin-zhu/structRFM/releases/latest/download/structRFM_checkpoint.tar.gz
tar -xzf structRFM_checkpoint.tar.gz
structRFM_checkpoint.export structRFM_checkpoint=PATH_TO_CHECKPOINT # modify ~/.bashrc for permanent setting
import os
from structRFM.infer import structRFM_infer
from_pretrained = os.getenv('structRFM_checkpoint')
model_paras = dict(max_length=514, dim=768, layer=12, num_attention_heads=12)
model = structRFM_infer(from_pretrained=from_pretrained, **model_paras)
seq = 'AGUACGUAGUA'
print('seq len:', len(seq))
feat_dic = model.extract_feature(seq)
for k, v in feat_dic.items():
print(k, v.shape)
'''
seq len: 11
cls_feat torch.Size([768])
seq_feat torch.Size([11, 768])
mat_feat torch.Size([11, 11])
'''
import os
from structRFM.model import get_structRFM
from structRFM.data import preprocess_and_load_dataset, get_mlm_tokenizer
from_pretrained = os.getenv('structRFM_checkpoint')
tokenizer = get_mlm_tokenizer(max_length=514)
model = get_structRFM(dim=768, layer=12, num_attention_heads=12, from_pretrained=from_pretrained, pretrained_length=None, max_length=514, tokenizer=tokenizer)
Requirements
Installation 0. Clone GitHub repo.
git clone https://github.com/heqin-zhu/structRFM.git
cd structRFM
conda env create -f structRFM_environment.yaml
conda activate structRFM
wget https://github.com/heqin-zhu/structRFM/releases/latest/download/structRFM_checkpoint.tar.gz
tar -xzf structRFM_checkpoint.tar.gz
structRFM_checkpoint.export structRFM_checkpoint=PATH_TO_CHECKPOINT # modify ~/.bashrc for permanent setting
The pretrianing sequence-structure dataset is constructed using RNAcentral and BPfold. We filter sequences with a length limited to 512, resulting about 21 millions sequence-structure paired data. It can be downloaded at Zenodo (4.5 GB).
USER_DIR and PROGRAM_DIR in scripts/run.sh,DATA_PATH and run_name in the following command,Then run:
bash scripts/run.sh --batch_size 96 --epoch 100 --lr 0.0001 --tag mlm --mlm_structure --max_length 514 --model_scale base --data_path DATA_PATH --run_name structRFM_512
For more information, run python3 main.py -h.
Download all data (3.7 GB) and checkpoints (2.2 GB) from Zenodo, and then place them into corresponding folder of each task.
We appreciate the following open-source projects for their valuable contributions:
If you find our work helpful, please cite our paper:
@article {structRFM,
author = {Zhu, Heqin and Li, Ruifeng and Zhang, Feng and Tang, Fenghe and Ye, Tong and Li, Xin and Gu, Yujie and Xiong, Peng and Zhou, S Kevin},
title = {A fully-open structure-guided RNA foundation model for robust structural and functional inference},
elocation-id = {2025.08.06.668731},
year = {2025},
doi = {10.1101/2025.08.06.668731},
publisher = {Cold Spring Harbor Laboratory},
URL = {https://www.biorxiv.org/content/early/2025/08/07/2025.08.06.668731},
journal = {bioRxiv}
}