Downloads · 30 days
5
15% of all-time downloads
Mhammad2023/my-dummy-model
my-dummy-model is a machine learning model from Mhammad2023. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
--- language: fr license: apache-2.0 tags: - masked-lm - camembert - transformers - tf - french - fill-mask ---
Downloads · 30 days
5
15% of all-time downloads
All-time downloads
34
Public
Repo size
544 MB
Likes
0
Public
Click a slice to open those files.
.h5543 MB · 99%
From the Hugging Face model README
language: fr license: apache-2.0 tags:
This is a TensorFlow-based masked language model (MLM) based on the camembert-base checkpoint, a RoBERTa-like model trained on French text.
This model uses the CamemBERT architecture, which is a RoBERTa-based transformer trained on large-scale French corpora (e.g., OSCAR, CCNet). It's designed to perform Masked Language Modeling (MLM) tasks.
It was loaded and saved using the transformers library in TensorFlow (TFAutoModelForMaskedLM). It can be used for fill-in-the-blank tasks in French.
from transformers import TFAutoModelForMaskedLM, AutoTokenizer
import tensorflow as tf
model = TFAutoModelForMaskedLM.from_pretrained("Mhammad2023/my-dummy-model")
tokenizer = AutoTokenizer.from_pretrained("Mhammad2023/my-dummy-model")
inputs = tokenizer("J'aime le [MASK] rouge.", return_tensors="tf")
outputs = model(**inputs)
logits = outputs.logits
masked_index = tf.argmax(inputs.input_ids == tokenizer.mask_token_id, axis=1)[0]
predicted_token_id = tf.argmax(logits[0, masked_index])
predicted_token = tokenizer.decode([predicted_token_id])
print(f"Predicted word: {predicted_token}")
This model inherits the limitations and biases from the camembert-base checkpoint, including:
Potential biases from the training data (e.g., internet corpora)
Use with caution in production or sensitive applications.
The model was not further fine-tuned; it is based directly on camembert-base, which was trained on:
OSCAR (Open Super-large Crawled ALMAnaCH coRpus)
CCNet (Common Crawl News)
No additional training was applied for this version. You can load and fine-tune it on your task using Trainer or Keras API.
This version has not been evaluated on downstream tasks. For evaluation metrics and benchmarks, refer to the original camembert-base model card.