Downloads · 30 days
6
12% of all-time downloads
cservan/multilingual-modernbert-large-diversity
multilingual-modernbert-large-diversity is a machine learning model from cservan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Pretrained multilingual language model using a masked language modeling (MLM) objective.
Downloads · 30 days
6
12% of all-time downloads
All-time downloads
50
Public
Repo size
3.8 GB
Likes
0
Public
Click a slice to open those files.
.bin1.9 GB · 100%
From the Hugging Face model README
Pretrained multilingual language model using a masked language modeling (MLM) objective.
ModernBERT is a transformers model pretrained on 1.2 billions of multilingual Wikipedia and OPUS tokens in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process to generate inputs and labels from those texts.
This model has the following configuration:
You can use the raw model for either masked language modeling or next sentence prediction, but it's mostly intended to be fine-tuned on a downstream task.
Note that this model is primarily aimed at being fine-tuned on tasks that use the whole sentence (potentially masked) to make decisions, such as sequence classification, token classification or question answering. For tasks such as text generation you should look at model like GPT ones.
Here is how to use this model to get the features of a given text in PyTorch:
from transformers import ModernBERTTokenizer, ModernBERTModel
tokenizer = ModernBERTTokenizer.from_pretrained('cservan/multilingual-modernbert-small')
model = ModernBERTModel.from_pretrained("cservan/multilingual-modernbert-small")
text = "Replace me by the text you want."
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
The ModernBERT model was pretrained on 1.2 billion of token from Multilingual Wikipedia (excluding lists, tables and headers) and OPUS.
It have been extracted from a 3.2 billion token corpus using an entropy sampling method called "diversity sampling" (paper to come).
The texts are cased and tokenized using SentencePiece and a vocabulary size of 128,000 tokens plus 1,000 unused token for downstream adataption. The inputs of the model are then of the form:
[CLS] Sentence A [SEP] Sentence B [SEP]
The tools used to pre-train the model are available here