Downloads · 30 days
0
seqSight/mgxlens-v2
mgxlens-v2 is a machine learning model from seqSight. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A retrieval-based metagenomic taxonomic profiler. Given short DNA reads (~150 bp) from a sequencing run, it reports the estimated relative abundance of microbial taxa at the genus and species level.
Downloads · 30 days
0
Access
Public
Updated Jul 14, 2026
Repo size
5.3 GB
Likes
0
Public
Click a slice to open those files.
.faiss2.4 GB · 83%
From the Hugging Face model README
A retrieval-based metagenomic taxonomic profiler. Given short DNA reads (~150 bp) from a sequencing run, it reports the estimated relative abundance of microbial taxa at the genus and species level.
Reference coverage. The index covers 13,399 distinct species-level SGBs across 3,297 genera — the full MetaPhlAn vJan25 SGB reference. Every indexed SGB resolves to a named Linnaean species (no unnamed placeholder bins among the indexed organisms; the naming rate over indexed SGBs is 100%). This means "closed-world" is about organisms outside MetaPhlAn's reference entirely — not gaps within it. An organism that is in MetaPhlAn's reference is representable; one that is not will be misassigned to its nearest reference genus (see Caveats → closed-world).
| File | What it is |
|---|---|
encoder.pt | The neural encoder (seqLens-89M backbone + attention pooling → 256-d embedding). Turns a DNA string into a vector. |
index.faiss | A FAISS vector index of ~2.32M reference marker embeddings. Cosine search via inner product. |
index.clades.npy | Row-aligned to the FAISS index: the SGB (clade) id for each indexed vector. |
index.markers.npy | Row-aligned marker source ids (provenance; not needed for inference). |
index.config.json | Index build parameters (window=150, nwin=5, emb_dim=256, max_length=128). |
lineage.tsv | clade_id → genus, species, display_name. Turns a retrieved SGB into a taxon name (display_name is MetaPhlAn-style). |
mgx_encoder.py | The encoder class + load_encoder() / embed() helpers. Import this; don't reimplement. |
MANIFEST.json | Machine-readable summary + file checksums + caveats. |
Encoder and index are from the same training run (index.config.json.source_ckpt ==
encoder.pt). Do not mix this encoder with any other index or vice-versa — the vectors would be
meaningless.
pip install torch transformers faiss-cpu numpy
# GPU optional: faiss-gpu + a CUDA torch build speed up large batches, but faiss-cpu is fine for
# moderate throughput. The encoder runs on CPU or GPU; GPU is ~10-50x faster for embedding.
The encoder downloads its base model (omicseye/seqLens_4096_512_89M-at-base-multi) from
HuggingFace on first load via transformers, so the serving host needs network access on first
run (or pre-cache the HF model). trust_remote_code=True is required for that base model.
For each input read:
index.clades.npy) of its top-1 hit. Map that clade
to genus/species via lineage.tsv.That's the whole pipeline. It's nearest-neighbor retrieval, not a trained classifier.
import numpy as np, faiss, torch
from collections import defaultdict
from mgx_encoder import load_encoder, embed # ships in this bundle
BUNDLE = "." # path to the bundle dir
K = 25 # neighbors per read
GATE = 0.66 # top-1 cosine threshold; below -> abstain
WINDOW = 150
# ---- load once at startup ----
device = "cuda" if torch.cuda.is_available() else "cpu"
model, tok, cfg = load_encoder(f"{BUNDLE}/encoder.pt", device)
index = faiss.read_index(f"{BUNDLE}/index.faiss")
clades = np.load(f"{BUNDLE}/index.clades.npy", allow_pickle=True) # (N,) SGB id per vector
lineage = {} # clade_id -> (genus, species, display_name)
with open(f"{BUNDLE}/lineage.tsv") as fh:
next(fh)
for line in fh:
cid, g, s, disp = line.rstrip("\n").split("\t")
lineage[cid] = (g, s, disp)
def center_window(seq, w=WINDOW):
if len(seq) <= w:
return seq
s = (len(seq) - w) // 2
return seq[s:s+w]
def profile(reads, batch_size=4096):
"""reads: list of DNA strings. Returns genus/species abundance + unclassified fraction.
Species keys are MetaPhlAn-style display names ('Escherichia coli', or
'Escherichia sp. (SGB123)' for unnamed genome bins)."""
genus_counts, species_counts = defaultdict(float), defaultdict(float)
n_total = len(reads); n_assigned = 0
for i in range(0, n_total, batch_size):
batch = [center_window(r) for r in reads[i:i+batch_size]]
ml = cfg["max_length"] if "max_length" in cfg else 128
Z = embed(model, tok, batch, device, max_length=ml).astype("float32")
D, I = index.search(Z, K) # D=cosine sims, I=index rows
for r in range(len(batch)):
if D[r, 0] < GATE: # gate: abstain
continue
clade = str(clades[I[r, 0]]) # top-1 hit's clade
g, s, disp = lineage.get(clade, ("", "", ""))
if not g:
continue
n_assigned += 1
genus_counts[g] += 1
if disp:
species_counts[disp] += 1 # display name, not raw s__ / SGB id
def norm(d):
tot = sum(d.values()) or 1.0
return {k: v/tot for k, v in sorted(d.items(), key=lambda x: -x[1])}
return {
"genus": norm(genus_counts),
"species": norm(species_counts),
"n_reads": n_total,
"n_assigned": n_assigned,
"unclassified_fraction": 1.0 - (n_assigned / n_total if n_total else 0.0),
}
# ---- example ----
if __name__ == "__main__":
reads = ["ACGT..."] # your reads (from a FASTQ)
print(profile(reads))
Reading FASTQ: reads come from .fastq/.fastq.gz. Use pysam, Bio.SeqIO, or a
2nd/4th-line parser to get the sequence strings, then pass the list to profile(). Paired-end
R1/R2 can both be fed as independent reads.
from fastapi import FastAPI, UploadFile
app = FastAPI()
# model/index/lineage loaded once at module import (see section 4)
@app.post("/profile")
async def profile_endpoint(file: UploadFile):
reads = parse_fastq(await file.read()) # your FASTQ parser -> list[str]
return profile(reads) # JSON: {genus:{...}, species:{...}, ...}
Load the model, index, and lineage once at startup (they're large — encoder 342 MB, index 2.3 GB). Never reload per request. The index holds ~2.3 GB in RAM; size the host accordingly (≥8 GB RAM recommended, ≥16 GB comfortable).
Input: DNA reads as strings (from FASTQ), ~150 bp each. Non-ACGT characters are tolerated by the tokenizer but degrade the embedding; upstream QC/trimming (e.g. fastp) is assumed.
Output (JSON):
{
"genus": {"escherichia": 0.29, "pseudomonas": 0.24, "...": 0.0},
"species": {"Escherichia coli": 0.21, "Pseudomonas aeruginosa": 0.18, "...": 0.0},
"n_reads": 200000,
"n_assigned": 94459,
"unclassified_fraction": 0.528
}
Abundances are fractions summing to ~1.0 within each level (over assigned reads). All 13,399
indexed SGBs have real species names, so species keys are Linnaean names (the Genus sp. (SGBxxxx)
fallback only appears if the index is rebuilt to include unnamed bins).
| Knob | Default | Effect |
|---|---|---|
GATE (top-1 cosine) | 0.66 | Higher = stricter, more reads abstained, fewer false assignments (but see Caveats — it does not reliably reject novel organisms). |
K (neighbors) | 25 | Only top-1 is used for assignment here; k>1 matters if you switch to k-NN vote aggregation. |
| aggregation | top-1 vote | Simplest. Alternatives (k-NN vote, similarity-weighted) exist in the research code but top-1 is the documented default. |
display_name logic still falls back to
Genus sp. (SGBxxxx) for unnamed bins — this only matters if the index is ever rebuilt to
include the unnamed SGBs that exist elsewhere in the MetaPhlAn taxonomy; the shipped index has
none.)enc_attention_bg1.best_genus.pt — seqLens-89M backbone, attention pooling, 256-d
projection, trained with contrastive InfoNCE + background negatives, checkpoint selected on
real-Zymo genus accuracy.index_bakeoff_mw_notest — leakage-fixed (held-out test markers excluded).omicseye/seqLens_4096_512_89M-at-base-multi.See MANIFEST.json for file checksums and the machine-readable summary.