Downloads · 30 days
17
23% of all-time downloads
alycialee/m2-bert-110M
m2-bert-110M is a fill-mask model from alycialee. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
The 110M checkpoint for M2-BERT-base from the paper Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture.
Downloads · 30 days
17
23% of all-time downloads
All-time downloads
75
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.bin463 MB · 100%
From the Hugging Face model README
The 110M checkpoint for M2-BERT-base from the paper Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture.
Check out our GitHub for instructions on how to download and fine-tune it!
You can load this model using Hugging Face AutoModel:
from transformers import AutoModelForMaskedLM
mlm = AutoModelForMaskedLM.from_pretrained('alycialee/m2-bert-110M', trust_remote_code=True)
This model uses the Hugging Face bert-base-uncased tokenizer:
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
You can use this model with a pipeline for masked language modeling:
from transformers import AutoModelForMaskedLM, BertTokenizer, pipeline
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
mlm = AutoModelForMaskedLM.from_pretrained('alycialee/m2-bert-110M', trust_remote_code=True)
unmasker = pipeline('fill-mask', model=mlm, tokenizer=tokenizer)
unmasker('Every morning, I enjoy a cup of [MASK] to start my day.')
This model requires trust_remote_code=True to be passed to the from_pretrained method. This is because we use custom PyTorch code (see our GitHub). You should consider passing a revision argument that specifies the exact git commit of the code, for example:
mlm = AutoModelForMaskedLM.from_pretrained(
'alycialee/m2-bert-110M',
trust_remote_code=True,
revision='0405c12',
)
Note use_flash_mm is false by default. Using FlashMM is currently not supported.
Using hyena_training_additions is turned off.