Downloads · 30 days
0
Shenzhi-Chen/DeepSTARR-Mouse
DeepSTARR-Mouse is a machine learning model from Shenzhi-Chen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for keras. The card lists the license as mit.
Model Summary DeepSTARR‑Mouse is a Convolutional Neural Network (CNN) adapted from the previously published DeepSTARR architecture (Nature Genetics, 2022). This model is designed for use in a transfer‑learning framewo…
Downloads · 30 days
0
Access
Public
Updated May 4, 2026
Repo size
277 MB
Likes
0
Public
Click a slice to open those files.
.h5106 MB · 98%
From the Hugging Face model README
Model Summary
DeepSTARR‑Mouse is a Convolutional Neural Network (CNN) adapted from the previously published DeepSTARR architecture (Nature Genetics, 2022).
This model is designed for use in a transfer‑learning framework to predict enhancer activity in E11.5 mouse embryos.
For each tissue, CNNs are pre‑trained on DNA accessibility data (i.e., ATAC‑seq) and fine‑tuned on a limited set of experimentally validated enhancers (VISTA enhancer browser, https://enhancer.lbl.gov/vista/).
Model Architecture
The DeepSTARR‑Mouse model is a custom TensorFlow 2 (Keras) implementation consisting of convolutional, pooling, and dense layers optimized to extract predictive features from 1,001‑bp DNA sequences.
The code and architecture definitions are available on GitHub:
➡️ https://github.com/Shenzhi‑Chen/DeepSTARR2‑Mouse
Model Weights
Sequence‑to‑accessibility and sequence‑to‑activity model weights are stored separately in two folders.
Each folder contains six tissue‑specific subfolders: heart, limb, midbrain (CNS), forebrain, hindbrain and neural tube
Within each tissue folder there are six models, corresponding to three cross‑validation folds and two replicates.
Model weights are stored in Model.json (architecture) and Model.h5 (trained parameters).
K27Ac_K4me1_models, promoter_distal_no_CTCF_models and tissueSpecATAC extended VISTA replacement transfer learning model weights are stored separately in three folders.
Each folder contains three tissue‑specific subfolders: heart, limb, midbrain
Within each tissue folder there are six models, corresponding to three cross‑validation folds and two replicates.
Model weights are stored in Model.json (architecture) and Model.h5 (trained parameters).
Training Objectives and Evaluation Metrics
Accessibility Model (Regression)
Activity Model (Classification)
Model Performance
Accessibility Model (Regression)
PCC values were computed between predicted and observed accessibility profiles across the held‑out test chromosome 18.
| Tissue | PCC |
|---|---|
| Heart | 0.76 |
| Limb | 0.76 |
| Midbrain | 0.78 |
Activity Model (Classification)
PPV values were computed on the held‑out test set using independent folds.
| Tissue | PPV (%) |
|---|---|
| Heart | 71.5 |
| Limb | 70.6 |
| Midbrain (CNS) | 80.2 |
All metrics represent the mean precision score across all cross‑validation folds and replicated models.
Dataset
Training and evaluation data are available in the companion Hugging Face dataset repository:
👉 Shenzhi‑Chen/DeepSTARR_Mouse_training_dataset
Framework
Implemented in TensorFlow 2 (v.2.4.1) / Keras.
All scripts for training, and using are hosted in the GitHub repository linked above.
Intended Use
This model is released for research and educational purposes.
It can be applied to predict enhancer activity or chromatin accessibility from DNA sequences in mammalian genomes, particularly for E11.5 mouse embryonic tissues.
Limitations
License
Citation
If you use this model, please cite:
[Paper citation once published]