Downloads · 30 days
18
10% of all-time downloads
GustavoHCruz/ExInDNABERT2
ExInDNABERT2 is a machine learning model from GustavoHCruz. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
DNABERT2 finetuned model for classifying DNA sequences into introns and exons, trained on a large cross-species GenBank dataset (34,627 different species).
Downloads · 30 days
18
10% of all-time downloads
All-time downloads
181
Public
Parameters
117M
468 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors468 MB · 100%
From the Hugging Face model README
DNABERT2 finetuned model for classifying DNA sequences into introns and exons, trained on a large cross-species GenBank dataset (34,627 different species).
You can use this model through its own custom pipeline:
from transformers import pipeline
pipe = pipeline(
task="dnabert2-exon-intron-classification",
model="GustavoHCruz/ExInDNABERT2",
trust_remote_code=True,
)
out = pipe(
"GCAGCAACAGTGCCCAGGGCTCTGATGAGTCTCTCATCACTTGTAAAG"
)
print(out) # EXON
This model uses the same maximum context length as the standard DNABERT2 (512 tokens), but it was trained on DNA sequences of up to 256 nucleotides.
The pipeline will automatically truncate the nucleotide sequence they exceed this limit.
The model expects the same tokens as DNABERT2, ou seja, nucleotídeos de entrada, como por exemplo
GTAAGGAGGGGGAT
The model should predict the class label: 0 (Intron) or 1 (Exon).
The model was trained on a processed version of GenBank sequences spanning multiple species, available at the DNA Coding Regions Dataset.
Average accuracy: 0.9956
| Class | Precision | Recall | F1-Score |
|---|---|---|---|
| Intron | 0.9943 | 0.9922 | 0.9932 |
| Exon | 0.9962 | 0.9972 | 0.9967 |
The full code for data processing, model training, and inference is available on GitHub:
CodingDNATransformers
You can find scripts for: