Downloads · 30 days
12
18% of all-time downloads
GustavoHCruz/NuclDNABERT2
NuclDNABERT2 is a machine learning model from GustavoHCruz. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
BERT finetuned model for classifying nucleotides into introns and exons, trained on a large cross-species GenBank dataset (34,627 different species).
Downloads · 30 days
12
18% of all-time downloads
All-time downloads
65
Public
Parameters
117M
468 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors468 MB · 100%
From the Hugging Face model README
BERT finetuned model for classifying nucleotides into introns and exons, trained on a large cross-species GenBank dataset (34,627 different species).
You can use this model through its own custom pipeline:
from transformers import pipeline
pipe = pipeline(
task="dnabert2-nucleotide-classification",
model="GustavoHCruz/NuclDNABERT2",
trust_remote_code=True,
)
out = pipe(
{
"before": "ATGATCCAGTTAAAAAATATATTC",
"sequence": "C",
"after": ""
}
)
print(out) # EXON
out = pipe(
{
"before": "GTAACATTAAAATAAAAAACAAAA",
"sequence": "T",
"after": "ATTATTTAAAGAAAAATATAATTA"
}
)
print(out) # INTRON
The maximum context of this model is the same as DNABERT2 (512 tokens), but its training was limited to 24 nucleotides before and after the target nucleotide.
When using the pipeline, these rules will be applied. The sequence will be limited to only a single nucleotide, and the sequences before and after will be truncated to a maximum of 24 nucleotides, even if longer sequences are provided. Similarly, the organism will be limited to 10 characters.
Prompt format:
The model expects the following input format:
TCGG...[SEP]T[SEP]AGCT...
The model should predict a class label: 0 (Exon), 1 (Intron) or 2 (Unknown).
The model was trained on a processed version of GenBank sequences spanning multiple species, available at the DNA Coding Regions Dataset.
Average accuracy: 0.7890
| Class | Precision | Recall | F1-Score |
|---|---|---|---|
| Intron | 0.6007 | 0.9113 | 0.6060 |
| Exon | 0.8011 | 0.9011 | 0.8482 |
| Unknown | 0.8096 | 0.6697 | 0.7331 |
The full code for data processing, model training, and inference is available on GitHub:
CodingDNATransformers
You can find scripts for: