Downloads · 30 days
0
LudensZhang/MGM2
MGM2 is a feature extraction model from LudensZhang. Use it when you need embeddings to search or compare text. It is set up for pytorch.
MGM2 converts a microbial community into contextualized microbial-token representations and one fixed-size community embedding. It combines 1,536-dimensional NTv3 microbial identity vectors with relative-abundance sig…
Downloads · 30 days
0
Access
Public
Updated Jul 15, 2026
Repo size
1.7 GB
Likes
1
Public
Click a slice to open those files.
.h51.4 GB · 84%
From the Hugging Face model README
MGM2 converts a microbial community into contextualized microbial-token representations and one fixed-size community embedding. It combines 1,536-dimensional NTv3 microbial identity vectors with relative-abundance signals in an abundance-aware Transformer.
| Directory | Hidden size | Layers | Heads | Community embedding |
|---|---|---|---|---|
models/small | 224 | 6 | 4 | 224 |
models/medium | 320 | 6 | 5 | 320 |
models/large | 384 | 8 | 6 | 384 |
models/xlarge | 512 | 10 | 8 | 512 |
Each model directory contains:
model.ckpt: an inference-only Lightning checkpoint derived from the
released last.ckpt, without optimizer state or training hyperparameters;config.json: the matching architecture and abundance configuration.Shared identity resources are stored once under embeddings/:
reference_ntv3_embeddings.h5: NTv3 embeddings keyed by reference OTU ID;taxonomy_fallback.h5: rank-aware taxonomy prototypes for inputs without
representative sequences.from huggingface_hub import snapshot_download
snapshot_download(
repo_id="<HUGGINGFACE_REPO_ID>",
local_dir="MGM2-release",
)
Replace <HUGGINGFACE_REPO_ID> with the ID of this model repository.
The inference utilities are maintained in the MGM2 source repository.
The FASTA record IDs must match the feature IDs in the abundance table.
python ntv3_embedding.py \
--fasta_path communities.fasta \
--model_path /path/to/NTv3_650M_pre \
--output_path community_ntv3_embeddings.h5 \
--gpu_ids 0
python src/data/build_dataset.py \
--otu_table community_table.csv \
--format csv \
--ntv3_embedding community_ntv3_embeddings.h5 \
--top_k_otus 768 \
--output community_input.pkl
Species names can be bare names such as Bacteroides fragilis, rank-prefixed
names such as s__Bacteroides fragilis, or complete taxonomy paths.
python src/data/build_dataset.py \
--otu_table community_species.csv \
--format csv \
--ntv3_embedding MGM2-release/embeddings/reference_ntv3_embeddings.h5 \
--fallback_embedding MGM2-release/embeddings/taxonomy_fallback.h5 \
--resolved_embedding community_resolved_embeddings.h5 \
--top_k_otus 768 \
--output community_input.pkl
Name matching is rank aware and uses only unambiguous names. The resolved HDF5
records embedding_source and matched_taxon_key for inspection.
python scripts/extract_community_embeddings.py \
--model_dir MGM2-release/models/small \
--data_file community_input.pkl \
--output community_embeddings.h5 \
--embedding_strategy cls \
--device cuda:0
cls is the recommended general-purpose representation. mean_pool exports
the mean of contextualized microbial tokens. qwen3_proj is available because
the released checkpoints include the alignment head.
HDF5 output contains sample_ids and an embeddings matrix with shape
[number_of_communities, hidden_size]. The extractor also supports NPZ and
CSV output by changing the output filename extension.
If you use MGM2 in your research, please cite the accompanying paper. Complete citation metadata will be added with the model release.