Downloads · 30 days
338
7% of all-time downloads
BSC-LT/roberta_model_for_anonimization
roberta_model_for_anonimization is a token classification model from BSC-LT. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This is a Roberta multilingual (Catalan & Spanish) anonimization model, for use with BSC's AnonymizationPipeline at:
Downloads · 30 days
338
7% of all-time downloads
All-time downloads
4.8K
Public
Repo size
993 MB
Likes
1
Public
Click a slice to open those files.
.bin496 MB · 99%
From the Hugging Face model README
This is a Roberta multilingual (Catalan & Spanish) anonimization model, for use with BSC's AnonymizationPipeline at:
https://github.com/TeMU-BSC/AnonymizationPipeline.
The anonymization pipeline is a library for performing sensitive data identification and ultimately anonymization of the detected data in Spanish and Catalan user generated plain text.
This is model can be used as a standalone model but it is meant to work within the pipeline.
The Roberta model can detect the following entities: ORG, PER, LOC
| Type | Score |
|---|---|
ENTS_F | 90.03 |
ENTS_P | 89.7 |
ENTS_R | 90.3 |