Downloads · 30 days
68
0% of all-time downloads
bertin-project/bertin-base-ner-conll2002-es
bertin-base-ner-conll2002-es is a token classification model from bertin-project. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as cc-by-4.0.
This checkpoint has been trained for the NER task using the CoNLL2002-es dataset.
Downloads · 30 days
68
0% of all-time downloads
All-time downloads
55.6K
Public
Parameters
124M
993 MB on disk
Likes
1
Public
Click a slice to open those files.
.bin496 MB · 50%
How the weights are stored.
F32124M · 100%
From the Hugging Face model README
This checkpoint has been trained for the NER task using the CoNLL2002-es dataset.
This is a NER checkpoint created from Bertin Gaussian 512, which is a RoBERTa-base model trained from scratch in Spanish. Information on this base model may be found at its own card and at deeper detail on the main project card.
The training dataset for the base model is mc4 subsampling documents to a total of about 50 million examples. Sampling is biased towards average perplexity values (using a Gaussian function), discarding more often documents with very large values (poor quality) of very small values (short, repetitive texts).
This is part of the Flax/Jax Community Week, organised by HuggingFace and TPU usage sponsored by Google.