Downloads · 30 days
24
8% of all-time downloads
biohub/scldm_cd4
scldm_cd4 is a machine learning model from biohub. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
scLDM.CD4 is a deep generative framework optimized for CD4+ T cell single-cell transcriptomics composed of an autoencoder and a flow-matching module. The autoencoder learns a rich latent representation of cell state,…
Downloads · 30 days
24
8% of all-time downloads
All-time downloads
311
Public
Parameters
208M
830 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors830 MB · 100%
From the Hugging Face model README
scLDM.CD4 is a deep generative framework optimized for CD4+ T cell single-cell transcriptomics composed of an autoencoder and a flow-matching module. The autoencoder learns a rich latent representation of cell state, and the flow model generates perturbed expression profiles conditioned on perturbation identity and context variables, with classifier-free guidance.
For Model Code and additional information on installation/usage please see the associated GitHub repository
scLDM.CD4 v0.1 builds on scLDM (Palla et al., 2025) and is tuned for counterfactual predictions of perturbation effects in single-cell transcriptomic profiles. The model has two main components:
Autoencoder: A transformer-based autoencoder for single-cell gene expression that compresses mRNA counts into a latent representation and reconstructs gene-level expression from that latent space.
Flow Matching (latent generative model): A conditional flow-matching model based on a Diffusion Transformer (DiT, Peebles et al., 2022) that generates latent cell profiles conditioned on attributes such as cell context and perturbation identity. Training uses an optimal-transport formulation to define target couplings along the flow starting from random noise, rather than constructing explicit couplings between control and perturbed cells.
Both architectures are built on PyTorch Lightning and support distributed training and inference. The models can be compiled with torch.compile for faster inference.
Parameter counts are architecture-dependent and configured through the training configuration. The released pre-trained checkpoint includes a 15.2M-parameter autoencoder and a 44.3M-parameter flow-matching model. The model also supports flexible architecture configurations, including:
Payam Dibaeinia (Biohub), Mei Knudson (Biohub, University of Chicago), Sudarshan Babu (Biohub), Jason Perera (Biohub), Aly A. Khan (Biohub, University of Chicago)
Payam Dibaeinia [email protected]
To submit feature requests or report issues with the model, please open an issue on the GitHub repository.
scLDM.CD4 v0.1 is designed to synthesize single-cell mRNA expression profiles of CD4+ T cells under single-gene knockdown perturbations, enabling downstream analysis of perturbation effects and counterfactual prediction. Key use cases include:
Do not use the model for the following purposes:
scLDM.CD4 v0.1 is trained on a pre-processed CD4+ T cell Perturb-seq dataset comprising ~14.5M cells, derived from the raw data released by Zhu et al. (2025). Training and evaluation are performed on a fixed panel of 3,699 highly variable genes (HVGs).
The training data include:
Training follows a two-stage approach:
Autoencoder Training: The autoencoder is trained first to learn a latent representation of gene expression data. Training includes:
Flow Matching Training: After autoencoder training, the flow matching model is trained in the learned latent space:
Training for both the autoencoder and the flow-matching model is launched via the shared experiments/scripts/train.py entry point and configured through a Hydra config:
experiments/scripts/train.py --config-name=marson_vaeexperiments/scripts/train.py --config-name=marson_fmTraining hyperparameters are configured via Hydra, with separate config files defined for the two stages.
Key hyper-parameters for training autoencoder include:
Key hyper-parameters for training flow-matching include:
A dedicated validation split of ~1.1M pre-processed cells is used to tune hyperparameters for both the autoencoder and flow-matching models (the final selected settings are recorded in the released marson_vae and marson_fm Hydra config files). Final results are reported on a held-out test split of ~2.4M pre-processed cells, disjoint from the training and validation data.
Optimal hyper-parameters were selected based on the following three key performance metrics:
On the held-out test split, we evaluate both generation quality and perturbation-effect prediction using: (1) UMAP comparisons of real vs. generated cells, (2) distributional metrics including MMD and W2, (3) Δ-based metrics that assess how well predicted mean shifts relative to the control population match true observed shifts, and (4) recovery of differentially expressed genes. Comparisons against simple but competitively strong baselines indicate that the model captures key aspects of the perturbation response.
The model may reflect biases present in the training data, including:
Areas of risk may include but are not limited to:
We are committed to advancing the responsible development and use of artificial intelligence. Please follow our Acceptable Use Policy when using the model.
Should you have any security or privacy issues or questions related to the model, please reach out to our team at [email protected] or [email protected], respectively.
We are deeply grateful to Giovanni Palla, Jakub Tomczak, Ali ElSheikh, Yibo Wen, Dennis Wu, Weimin Wu, Krunal Patel, Lakshmi Krishnan, Shirin Fuller, and Kavita Kulkarni for their generous support, thoughtful feedback, and many helpful discussions throughout this work.