Downloads · 30 days
21
14% of all-time downloads
Jonandrop/camelbert-lemma-ar
camelbert-lemma-ar is a token classification model from Jonandrop. Use it when you need labels on individual words, such as names. It is set up for onnx. The card lists the license as apache-2.0.
Per-language lemmatizer and UPOS tagger fine-tuned from CAMeL-Lab/bert-base-arabic-camelbert-msa on UD Arabic-PADT. Predicts an edit-tree label per token; applying it to the surface word yields the lemma. Dynamic-int8…
Downloads · 30 days
21
14% of all-time downloads
All-time downloads
145
Public
Repo size
234 MB
Likes
0
Public
Click a slice to open those files.
.onnx117 MB · 96%
From the Hugging Face model README
Per-language lemmatizer and UPOS tagger fine-tuned from CAMeL-Lab/bert-base-arabic-camelbert-msa on UD Arabic-PADT. Predicts an edit-tree label per token; applying it to the surface word yields the lemma. Dynamic-int8 quantized ONNX (~112 MB).
| Metric | int8 | fp (reference) |
|---|---|---|
| Lemma accuracy | 43.91% | 43.91% |
| UPOS accuracy | 68.32% | N/A |
int8 trades accuracy for size; the fp16/torch model is several points higher (UPOS head is most affected by quantization).
Plain argmax over edit-tree labels often picks a label that is structurally invalid for the word (its delete-segment does not match), applying it fails and accuracy collapses. At inference you MUST iterate candidate labels in descending logit order and pick the first whose edit applies to the word (identity is the guaranteed fallback). This is model-internal selection, not a hand-written rule.
model.int8.onnx: inputs input_ids, attention_mask; outputs upos_logits, lemma_logitsconfig.json, tokenizer.json, tokenizer_config.json, special_tokens_map.jsonid2label.json / label2id.json: edit-tree lemma labelsupos_id2label.json / upos_label2id.json: UPOS labelsedit_trees.json: edit-tree label definitionslexicon.json: fallback {word: lemma} lexicon