Downloads · 30 days
0
blazerye/AROMA
AROMA is a machine learning model from blazerye. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<p align="center" <img src="figures/logo.jpg" alt="AROMA Logo" width="120" </p
Downloads · 30 days
0
Access
Public
Updated Apr 23, 2026
Parameters
8.5B
19.8 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors17.5 GB · 97%
How the weights are stored.
BF168.2B · 97%
From the Hugging Face model README
Please refer to our repository and paper for more details.
AROMA is a novel multimodal architecture for virtual cell modeling that integrates textual evidence, graph topology, and protein sequences to predict the effects of genetic perturbations.
<p align="center"> <img src="figures/overview.jpg" alt="Overview"> </p>The overall AROMA pipeline is illustrated in the figure above and is divided into three stages:
Data stage. AROMA constructs two complementary knowledge graphs and a large-scale virtual cell reasoning dataset for evidence grounding.
Modeling stage. AROMA adopts a retrieval-augmented strategy to incorporate query-relevant information, thereby providing explicit evidence cues for prediction. In addition, it jointly leverages topological representations learned from graph neural networks (GNN) and protein sequence representations encoded by ESM-2, and applies a cross-attention module to explicitly model perturbation-target gene dependencies across modalities.
Training stage. AROMA first performs multimodal supervised fine-tuning (SFT), and is then further optimized with Group Relative Policy Optimization (GRPO) reinforcement learning to enhance predictive performance while generating biologically meaningful explanations.