Downloads · 30 days
12
12% of all-time downloads
hicai-zju/InstructBioMol-base
InstructBioMol-base is a machine learning model from hicai-zju. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Downloads · 30 days
12
12% of all-time downloads
All-time downloads
97
Public
Repo size
13.5 GB
Likes
0
Public
Click a slice to open those files.
.bin13.5 GB · 100%
From the Hugging Face model README
InstructBioMol is a multimodal large language model that bridges natural language with biomolecules (proteins and small molecules). It achieves any-to-any alignment between natural language, molecules, and proteins through comprehensive instruction tuning.
For detailed information, please refer to our paper and code repository.
| Model Name | Stage | Multimodal | Description |
|---|---|---|---|
| InstructBioMol-base (This Model) | Pretraining | ❎ | Continual pretrained model on molecular sequences, protein sequences, and scientific literature. |
| InstructBioMol-instruct-stage1 | Instruction tuning (stage 1) | ✅ | Stage1 instruction-tuned model with biomolecular multimodal processing capabilities. (e.g., 3D molecules/proteins) |
| InstructBioMol-instruct | Instruction tuning (stage 1 and 2) | ✅ | Fully instruction-tuned model (stage1 & stage2) with biomolecular multimodal processing capabilities (e.g., 3D molecules/proteins) |
Base Architecture: LLaMA-2-7B
Training Data:
1. Molecular Sequences:
2. Protein Sequences:
<p> (e.g., <p>M<p>A<p>L<p>W...).3. Natural Language Texts:
Training Objective: Causal language modeling (self-supervised)
from transformers import LlamaForCausalLM, LlamaTokenizer
import torch
model_name = "hicai-zju/InstructBioMol-base"
tokenizer = LlamaTokenizer.from_pretrained(model_name)
model = LlamaForCausalLM.from_pretrained(model_name, device_map="cuda:0")
prompt = "<p>M" # protein sequence
# prompt = "[C]" # molecule sequence
# prompt = 'Scientific' # natural language
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.7,
top_p=0.9,
do_sample=True
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(generated_text)
@article{zhuang2025advancing,
author = {Xiang Zhuang and
Keyan Ding and
Tianwen Lyu and
Yinuo Jiang and
Xiaotong Li and
Zhuoyi Xiang and
Zeyuan Wang and
Ming Qin and
Kehua Feng and
Jike Wang and
Qiang Zhang and
Huajun Chen},
title={Advancing biomolecular understanding and design following human instructions},
journal={Nature Machine Intelligence},
pages={1--14},
year={2025},
publisher={Nature Publishing Group UK London}
}