Downloads · 30 days
24
3% of all-time downloads
hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1
M2-BERT-2k-Retrieval-Encoder-V1 is a fill-mask model from hazyresearch. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
The 80M checkpoint for M2-BERT-2k from the paper Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT.
Downloads · 30 days
24
3% of all-time downloads
All-time downloads
874
Public
Parameters
82M
328 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors328 MB · 100%
How the weights are stored.
F3282M · 100%
From the Hugging Face model README
The 80M checkpoint for M2-BERT-2k from the paper Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT.
Check out our GitHub for instructions on how to download and fine-tune it!
You can load this model using Hugging Face AutoModel:
from transformers import AutoModelForMaskedLM, BertConfig
config = BertConfig.from_pretrained("hazyresearch/M2-BERT-2K-Retrieval-Encoder-V1")
model = AutoModelForMaskedLM.from_pretrained("hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1", config=config, trust_remote_code=True)
This model uses the Hugging Face bert-base-uncased tokenizer:
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
This model generates embeddings for retrieval. The embeddings have a dimensionality of 768:
from transformers import AutoTokenizer, AutoModelForMaskedLM, BertConfig
max_seq_length = 2048
testing_string = "Every morning, I make a cup of coffee to start my day."
config = BertConfig.from_pretrained("hazyresearch/M2-BERT-2K-Retrieval-Encoder-V1")
model = AutoModelForMaskedLM.from_pretrained("hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1", config=config, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased", model_max_length=max_seq_length)
input_ids = tokenizer([testing_string], return_tensors="pt", padding="max_length", return_token_type_ids=False, truncation=True, max_length=max_seq_length)
outputs = model(**input_ids)
embeddings = outputs['sentence_embedding']
This model requires trust_remote_code=True to be passed to the from_pretrained method. This is because we use custom PyTorch code (see our GitHub). You should consider passing a revision argument that specifies the exact git commit of the code, for example:
mlm = AutoModelForMaskedLM.from_pretrained(
"hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1",
config=config,
trust_remote_code=True,
)
Note use_flash_mm is false by default. Using FlashMM is currently not supported.