Downloads · 30 days
18
20% of all-time downloads
Pangeanic/mapa-multilingual-administrative
mapa-multilingual-administrative is a token classification model from Pangeanic. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
This model is part of the MAPA (Multilingual Anonymisation for Public Administrations) toolkit, developed by Pangeanic and funded by the European Union through the Connecting Europe Facility (CEF) programme.
Downloads · 30 days
18
20% of all-time downloads
All-time downloads
91
Public
Repo size
1.4 GB
Likes
0
Public
Click a slice to open those files.
.bin714 MB · 100%
From the Hugging Face model README
This model is part of the MAPA (Multilingual Anonymisation for Public Administrations) toolkit, developed by Pangeanic and funded by the European Union through the Connecting Europe Facility (CEF) programme.
It performs hierarchical Named Entity Recognition (NER) for the detection of personally identifiable information (PII) in multilingual text, with the goal of supporting anonymisation workflows in public administrations.
google-bert/bert-base-multilingual-cased (mBERT cased), with extended vocabularyEnhancedTwoFlatLevelsSequenceLabellingModel — a custom BERT-based architecture with two parallel classification headsThe model performs token classification at two levels simultaneously:
PERSON, ORGANISATION, LOCATION, DATE, ADDRESS...).PERSON: title, given name, family name...).Example output structure:
{
"annotations": [
{ "content": "señor Connelly", "value": "PERSON" },
{ "content": "señor", "value": "title" },
{ "content": "Connelly", "value": "family name" }
]
}
The full label inventories are included in this repository as level1_tags_vocabulary.json and level2_tags_vocabulary.json.
This model is intended to be used as part of the MAPA toolkit, which provides the full anonymisation pipeline (entity detection + entity replacement).
The model uses a custom architecture (EnhancedTwoFlatLevelsSequenceLabellingModel) defined in the MAPA codebase. It is not directly loadable via AutoModel.from_pretrained from the transformers library without that code.
To use this model, clone the MAPA toolkit:
🔗 Repository: https://github.com/PangeanicAI/MAPA-EU-Project
The repository contains the model class definition, inference scripts, and the complete anonymisation pipeline (including the entity replacement module, which uses auxiliary resources not included in this HF repo).
Final metrics reported at the end of training:
| Metric | Value |
|---|---|
| Level 1 micro-F1 | 0.8374 |
| Level 1 binary-F1 | 0.8669 |
| Level 2 micro-F1 | 0.8467 |
| Final loss | 6.2376 |
Other models from the MAPA project are available under the Pangeanic organisation, covering additional languages and domains.
The MAPA project was funded by the European Union under the Connecting Europe Facility (CEF) programme, grant agreement INEA/CEF/ICT/A2019/1927065.
If you use this model, please cite the MAPA project:
@inproceedings{mapa2022,
title = {{MAPA} Project: Ready-to-Go Open-Source Datasets and Deep Learning Technology to Remove Identifying Information from Text Documents},
author = {Gianola, Lucia and Ajausks, \=Emils and Arranz, Victoria and Bendi, Chomicha and Choukri, Khalid and Ciulla, Montse and Coheur, Luísa and Costa, Costanza and Cruz, Elena and Esplà-Gomis, Miquel and Garcia-Martinez, Mercedes and Herranz, Manuel and Iranzo-Sánchez, Javier and Klūga, Mārcis and Labaka, Gorka and Lagzdiņš, Artūrs and Lazar, Alina and Mahdi, Mohammed and Otero, Carla and Pinnis, Mārcis and Rigau, German and Ryšavá, Klára and Saint-Dizier, Patrick and Sosoni, Vilelmini},
booktitle = {Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-10)},
year = {2022}
}
Read more -
For questions about the MAPA toolkit, please refer to the project repository or contact Pangeanic.