Downloads · 30 days
0
almanach/camembertv2-base-sequoia
camembertv2-base-sequoia is a token classification model from almanach. Use it when you need labels on individual words, such as names. The card lists the license as mit.
almanach/camembertv2-base-sequoia is a roberta model for token classification. It is trained on the Sequoia dataset for the task of Part-of-Speech Tagging and Dependency Parsing. The model achieves an f1 score of on t…
Downloads · 30 days
0
Access
Public
Updated Nov 14, 2024
Repo size
2.6 GB
Likes
0
Public
Click a slice to open those files.
.pt1.7 GB · 69%
From the Hugging Face model README
almanach/camembertv2-base-sequoia is a roberta model for token classification. It is trained on the Sequoia dataset for the task of Part-of-Speech Tagging and Dependency Parsing. The model achieves an f1 score of on the Sequoia dataset.
The model is part of the almanach/camembertv2-base family of model finetunes.
The model can be used for token classification tasks in French for Part-of-Speech Tagging and Dependency Parsing.
The model may exhibit biases based on the training data. The model may not generalize well to other datasets or tasks. The model may also have limitations in terms of the data it was trained on.
You can use the models directly with the hopsparser library in server mode https://github.com/hopsparser/hopsparser/blob/main/docs/server.md
Model trained with the hopsparser library on the Sequoia dataset.
# Layer dimensions
mlp_input: 1024
mlp_tag_hidden: 16
mlp_arc_hidden: 512
mlp_lab_hidden: 128
# Lexers
lexers:
- name: word_embeddings
type: words
embedding_size: 256
word_dropout: 0.5
- name: char_level_embeddings
type: chars_rnn
embedding_size: 64
lstm_output_size: 128
- name: fasttext
type: fasttext
- name: camembertv2_base_p2_17k_last_layer
type: bert
model: /scratch/camembertv2/runs/models/camembertv2-base-bf16/post/ckpt-p2-17000/pt/
layers: [11]
subwords_reduction: "mean"
# Training hyperparameters
encoder_dropout: 0.5
mlp_dropout: 0.5
batch_size: 8
epochs: 64
lr:
base: 0.00003
schedule:
shape: linear
warmup_steps: 100
UPOS: 0.99383 LAS: 0.94942
roberta custom model for token classification.
BibTeX:
@misc{antoun2024camembert20smarterfrench,
title={CamemBERT 2.0: A Smarter French Language Model Aged to Perfection},
author={Wissam Antoun and Francis Kulumba and Rian Touchent and Éric de la Clergerie and Benoît Sagot and Djamé Seddah},
year={2024},
eprint={2411.08868},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.08868},
}
@inproceedings{grobol:hal-03223424,
title = {Analyse en dépendances du français avec des plongements contextualisés},
author = {Grobol, Loïc and Crabbé, Benoît},
url = {https://hal.archives-ouvertes.fr/hal-03223424},
booktitle = {Actes de la 28ème Conférence sur le Traitement Automatique des Langues Naturelles},
eventtitle = {TALN-RÉCITAL 2021},
venue = {Lille, France},
pdf = {https://hal.archives-ouvertes.fr/hal-03223424/file/HOPS_final.pdf},
hal_id = {hal-03223424},
hal_version = {v1},
}