Downloads · 30 days
10
2% of all-time downloads
buetnlpbio/bidna-bert
bidna-bert is a fill-mask model from buetnlpbio. Use it when you need the model to fill a missing word. It is set up for transformers.
<span style="color:red"buetnlpbio/bidna-bert was trained on only Human Genome DNA dataset for 1 epoch (for ablation). Performance on other DNA types may be limited. </span
Downloads · 30 days
10
2% of all-time downloads
All-time downloads
424
Public
Repo size
937 MB
Likes
0
Public
Click a slice to open those files.
.bin468 MB · 100%
From the Hugging Face model README
<span style="color:red">buetnlpbio/bidna-bert was trained on only Human Genome DNA dataset for 1 epoch (for ablation). Performance on other DNA types may be limited. </span>
BiRNA-BERT is a BERT-style transformer encoder model that generates embeddings for RNA sequences. BiRNA-BERT has been trained on BPE tokens and individual nucleotides. As a result, it can generate both granular nucleotide-level embeddings and efficient sequence-level embeddings (using BPE).
BiRNA-BERT was trained using the MosaicBERT framework - https://huggingface.co/mosaicml/mosaic-bert-base
import torch
import transformers
from transformers import AutoModelForMaskedLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("buetnlpbio/bidna-tokenizer")
config = transformers.BertConfig.from_pretrained("buetnlpbio/bidna-bert")
mysterybert = AutoModelForMaskedLM.from_pretrained("buetnlpbio/bidna-bert",config=config,trust_remote_code=True)
mysterybert.cls = torch.nn.Identity()
# To get sequence embeddings
seq_embed = mysterybert(**tokenizer("AGCTACGTACGT", return_tensors="pt"))
print(seq_embed.logits.shape) # CLS + 4 BPE token embeddings + SEP
# To get nucleotide embeddings
char_embed = mysterybert(**tokenizer("A G C T A C G T A C G T", return_tensors="pt"))
print(char_embed.logits.shape) # CLS + 12 nucleotide token embeddings + SEP