Downloads · 30 days
0
emalek/RiboSignal
RiboSignal is a other model from emalek. Use it for the other task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Predicts a per-nucleotide ribosome P-site profile for a transcript from its sequence and matched RNA-seq coverage. It takes any transcript, coding or non-coding, up to 10,000 nt, the longest in training. No ribosome-p…
Downloads · 30 days
0
Access
Public
Updated Sep 27, 2026
Repo size
50.4 MB
Likes
0
Public
Click a slice to open those files.
.pt50.4 MB · 100%
From the Hugging Face model README
Predicts a per-nucleotide ribosome P-site profile for a transcript from its sequence and matched RNA-seq coverage. It takes any transcript, coding or non-coding, up to 10,000 nt, the longest in training. No ribosome-profiling experiment is required at inference.
The predicted profile can be fed to an ORF caller in place of real Ribo-seq, which is how it is evaluated here and what it is for: calling translated ORFs, including upstream and non-canonical ones, in samples where no Ribo-seq exists.
Two checkpoints are released. They differ only in the sequence mixer; body, heads, inputs, training data and ORF track are identical.
| checkpoint | mixer | params | device | profile r, 3-seed mean (range) |
|---|---|---|---|---|
mamba4_best.pt (primary) | dilated CNN + 4 bidirectional Mamba blocks | 7,521,026 | GPU; CPU through the repository's pure-PyTorch reference | 0.6799 (0.6753-0.6851) |
attn_best.pt | dilated CNN + 2 transformer layers | 5,071,106 | CPU or GPU | 0.6595 (0.6585-0.6603) |
Profile r is the median per-transcript Pearson correlation on raw counts over 70,883 held-out
transcripts, averaged across three training seeds. The released checkpoints are the seed-0 runs:
0.6851 for mamba4 and 0.6585 for attn, the values in <arch>_test_metrics.json.
Ten channels per nucleotide:
| channels | content |
|---|---|
| 4 | one-hot A/C/G/T of the transcript |
| 5 | ORF-candidate track: frame 0/1/2 occupancy, start context, stop |
| 1 | RNA-seq coverage, depth-normalised |
Two output heads: a per-nucleotide profile (a distribution summing to 1) and a per-transcript count.
The inputs are built from your own transcripts and RNA-seq with the repository's pack builder; the tutorial at https://ribosignal.readthedocs.io does this for chromosome 22. Two settings must match what the checkpoints were trained on:
--mode ext --kozak none. Both released models are
--kozak none models; feeding them a heuristic-Kozak track is a mistake that has already
happened once in this project.Reproducing the held-out metrics above additionally needs the training universe (84,472 transcripts), which is not published here.
import torch, json
cfg = json.load(open("mamba4_config.json"))
sd = torch.load("mamba4_best.pt", map_location="cpu") # plain state_dict
model.py is the model definition. With the copy here, mamba4_best.pt requires mamba_ssm
and a CUDA device. The repository's scripts/model.py falls back to a pure-PyTorch reference
(scripts/mamba_ref.py) where mamba_ssm is absent, so mamba4 also runs on CPU, more slowly.
attn_best.pt has no such requirement.
--outFilterMultimapNmax 1); rRNA, tRNA, miRNA and
mitochondrial loci removed before P-site calling. Mitochondrial protein-coding genes excluded
throughout.count_weight 0.1, input_mode: both.pearson_median.MIT. Free for any use including commercial, with no attribution requirement beyond keeping the
copyright notice. Full text in LICENSE.
The weights and model.py are covered. The training data are not: they are third-party public
datasets (GEO GSE182371 and GSE182372) under their own terms, and this licence makes no claim
about them.
https://github.com/ericmalekos/ribosignal
model.py here is an earlier copy of that repository's scripts/model.py; the repository version
adds the CPU fallback for mamba4. The pack builder, prediction, ORF-calling and scoring pipeline
live there.
Manuscript in preparation. Until then cite this repository and https://github.com/ericmalekos/ribosignal.