Downloads · 30 days
15
0% of all-time downloads
ERCDiDip/40_langdetect_v01
40_langdetect_v01 is a text classification model from ERCDiDip. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of xlm-roberta-base on the monasterium.net dataset.
Downloads · 30 days
15
0% of all-time downloads
All-time downloads
31.8K
Public
Repo size
4.5 GB
Likes
0
Public
Click a slice to open those files.
.bin1.1 GB · 98%
From the Hugging Face model README
This model is a fine-tuned version of xlm-roberta-base on the monasterium.net dataset.
On the top of this XLM-RoBERTa transformer model is a classification head. Please refer this model together with to the XLM-RoBERTa (base-sized model) card or the paper Unsupervised Cross-lingual Representation Learning at Scale by Conneau et al. for additional information.
You can directly use this model as a language detector, i.e. for sequence classification tasks. Currently, it supports the following 41 languages, modern and medieval:
Modern: Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), German (de), Greek (el), Hungarian (hu), Irish (ga), Italian (it), Latvian (lv), Lithuanian (lt), Maltese (mt), Polish (pl), Portuguese (pt), Romanian (ro), Slovak (sk), Slovenian (sl), Spanish (es), Swedish (sv), Russian (ru), Turkish (tr), Basque (eu), Catalan (ca), Albanian (sq), Serbian (se), Ukrainian (uk), Norwegian (no), Arabic (ar), Chinese (zh), Hebrew (he)
Medieval: Middle High German (mhd), Latin (la), Middle Low German (gml), Old French (fro), Old Church Slavonic (chu), Early New High German (fnhd), Ancient and Medieval Greek (grc)
The model was fine-tuned using the Monasterium and Wikipedia datasets, which consist of text sequences in 40 languages. The training set contains 80k samples, while the validation and test sets contain 16k. The average accuracy on the test set is 99.59% (this matches the average macro/weighted F1-score, the test set being perfectly balanced).
Fine-tuning was done via the Trainer API with WeightedLossTrainer.
The following hyperparameters were used during training:
mixed_precision_training: Native AMP
| Training Loss | Validation Loss | F1 |
|---|---|---|
| 0.000300 | 0.048985 | 0.991585 |
| 0.000100 | 0.033340 | 0.994663 |
| 0.000000 | 0.032938 | 0.995979 |
#Install packages
!pip install transformers --quiet
#Import libraries
import torch
from transformers import pipeline
#Define pipeline
classificator = pipeline("text-classification", model="ERCDiDip/40_langdetect_v01")
#Use pipeline
classificator("clemens etc dilecto filio scolastico ecclesie wetflari ensi treveren dioc salutem etc significarunt nobis dilecti filii commendator et fratres hospitalis beate marie theotonicorum")
Please cite the following papers when using this model.
@misc{ercdidip2022,
title={40 langdetect v01 (Revision 9fab42a)},
author={Kovács, Tamás, Atzenhofer-Baumgartner, Florian, Aoun, Sandy, Nicolaou, Anguelos, Luger, Daniel, Decker, Franziska, Lamminger, Florian and Vogeler, Georg},
year = { 2022 },
url = { https://huggingface.co/ERCDiDip/40_langdetect_v01 },
doi = { 10.57967/hf/0099 },
publisher = { Hugging Face }
}
This model is part of the From Digital to Distant Diplomatics (DiDip) ERC project funded by the European Research Council.