Downloads · 30 days
15
5% of all-time downloads
antoinelouis/camemberta-L2
camemberta-L2 is a feature extraction model from antoinelouis. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
This model is a pruned version of the pre-trained CamemBERTa checkpoint, obtained by dropping the top-layers from the original model.
Downloads · 30 days
15
5% of all-time downloads
All-time downloads
275
Public
Parameters
40.9M
165 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors164 MB · 98%
From the Hugging Face model README
This model is a pruned version of the pre-trained CamemBERTa checkpoint, obtained by dropping the top-layers from the original model.
You can use the raw model for masked language modeling (MLM), but it's mostly intended to be fine-tuned on a downstream task, especially one that uses the whole sentence to make decisions such as text classification, extractive question answering, or semantic search. For tasks such as text generation, you should look at autoregressive models like BelGPT-2.
You can use this model directly with a pipeline for masked language modeling:
from transformers import pipeline
unmasker = pipeline('fill-mask', model='antoinelouis/camemberta-L2')
unmasker("Bonjour, je suis un [MASK] modèle.")
You can also use this model to extract the features of a given text:
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained('antoinelouis/camemberta-L2')
model = AutoModel.from_pretrained('antoinelouis/camemberta-L2')
text = "Remplacez-moi par le texte de votre choix."
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
CamemBERTa has originally been released in a base (112M) version. The following checkpoints prune the base variation by dropping the top 2, 4, 6, 8, and 10 pretrained encoding layers, respectively.
| Model | #Params | Size | Pruning |
|---|---|---|---|
| CamemBERTa-base | 111.8M | 447MB | - |
| CamemBERTa-L10 | 97.6M | 386MB | -14% |
| CamemBERTa-L8 | 83.5M | 334MB | -25% |
| CamemBERTa-L6 | 69.3M | 277MB | -38% |
| CamemBERTa-L4 | 55.1M | 220MB | -51% |
| CamemBERTa-L2 | 40.9M | 164MB | -63% |