Downloads ยท 30 days
0
BGI-HangzhouAI/Gengram
Gengram is a machine learning model from BGI-HangzhouAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div align="center" style="line-height: 2;" <a href="https://huggingface.co/BGI-HangzhouAI/Gengram" target="blank" <img alt="Hugging Face" src="https://img.shields.io/badge/๐ค%20Hugging%20Face-Gengram%20-ffc107"/ </aโฆ
Downloads ยท 30 days
0
Access
Public
Updated Feb 3, 2026
Repo size
71.8 GB
Likes
2
Public
Click a slice to open those files.
.pt71.8 GB ยท 100%
From the Hugging Face model README
Gengram is a novel conditional memory module designed for genomic foundation models (GFMs) that introduces explicit motif memory retrieval to enhance Transformer-based DNA sequence modeling. Unlike traditional GFMs that rely on dense computation to implicitly infer multi-nucleotide motifs, Gengram provides an efficient lookup mechanism for biological patterns through a genomic-specific hashing scheme.
Figure 1 illustrates the overall architecture of Gengram, together with the evaluation pipeline used to assess its effectiveness across multiple genomic benchmarks.

Gengram exhibits clear biologically grounded behaviors, including:
The following details the model configuration, including the parameterization of Gengram, MoE routing strategies, and training hyperparameters used across all experiments.
Gengram Parameters
These parameters control how Gengram operates within the Transformer layers, including which layers to apply it to, the n-gram sizes, and embedding dimensions.
| Parameter | Description | Example |
|---|---|---|
--gengram-enabled | Enable Gengram | true |
--gengram-layer-ids | Layers to apply Gengram | 3 6 10 |
--gengram-ngram-sizes | N-gram sizes for DNA processing | 1 2 3 4 5 6 |
--gengram-embed-dim-per-ngram | Embedding dimension per n-gram | 1024 |
--gengram-window-size | window size | 21 |
Mixture of Experts (MoE)
These parameters define the Mixture-of-Experts architecture, including the number of experts, routing top-k, and load balancing strategies during training.
| Parameter | Description | Default |
|---|---|---|
--num-experts | Number of experts | 8 |
--moe-router-topk | Top-k experts to route to | 2 |
--moe-router-load-balancing-type | Load balancing strategy | aux_loss |
--moe-aux-loss-coeff | Auxiliary loss coefficient | 1e-3 |
Training Parameters
These parameters specify the training setup, including sequence length, batch sizes, precision, and attention optimizations.
| Parameter | Description | Example |
|---|---|---|
--seq-length | Maximum sequence length | 8192 |
--micro-batch-size | Micro batch size per GPU | 1 |
--global-batch-size | Global batch size across all GPUs | 1024 |
--bf16 | Use BF16 precision | true |
--use-flash-attn | Enable Flash Attention | true |
Gengram demonstrates strong performance across multiple genomic benchmarks, achieving competitive results despite being trained on significantly fewer tokens and with a smaller model size.
<div align="center">| Metric | Gengram-10B | Genos-10B | Evo2-40B |
|---|---|---|---|
| Trained Tokens | 200B | 2.2T | 9.3T |
| Multi-species Exon Classification | 0.9832 | 0.9755 | 0.9332 |
| Splice Site Identification | 0.9009 | 0.7990 | 0.9138 |
| Human OCR Ensembl | 0.7714 | 0.7623 | 0.7635 |
Key Observations
Evaluation Benchmarks
Gengram model is available for download from Hugging Face. We provide torch version.
| Model | Activated Params | Hugging Face | Format |
|---|---|---|---|
| Gengram-10B | 2.87 B | ๐ค Hugging Face | torch |
Run the pre-training script with the following command:
cd Gengram
bash Gengram_layer3-6-10_win21_pp2.sh
This repository and the Gengram model weights are licensed under the Apache License 2.0.
Please note that the primary use of Gengram model is to support genomics research, providing researchers with advanced analytical capabilities and long-context modeling tools powered by large-scale foundation models for the human genome. It is not intended for use in any manner that violates applicable laws or regulations, nor for any activities prohibited by the license agreement.
We acknowledge the high-quality sequencing data provided by CycloneSEQ, which forms an important foundation for this work. We also appreciate the inspiration from DeepSeek's Engram module and the framework support provided by Megatron-LM. Model training was conducted on the 021 Science Foundation Model and Zero2X open platform.
If you use this work in your research, please cite the following paper:
@article@article{gengram2026,
title={Beyond Conditional Computation: Retrieval-Augmented Genomic Foundation Models with Gengram},
author={Genos Team and Xu, Huinan and Feng, Xuyang and Chen, Junhong and Liu Junchen and Deng, Kaiwen and Ding, Kai and Long, Shengning and Shuai, Jiaxue and Li, Zhaorong and Liu, Shiping and Xue, Guirong and Xiao, Zhan},
journal={arXiv preprint arXiv:2601.22203},
year={2026}
}
For project-related questions, please open an issue. You can also contact the Genos Team at [email protected].