Downloads · 30 days
25.1K
8% of all-time downloads
DeepMount00/Italian_NER_XXL
Italian_NER_XXL is a token classification model from DeepMount00. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
25.1K
8% of all-time downloads
All-time downloads
325K
Public
Parameters
110M
3.5 GB on disk
Likes
50
Public
Click a slice to open those files.
.bin441 MB · 50%
From the Hugging Face model README
💡 Found this resource helpful? Creating and maintaining open source AI models and datasets requires significant computational resources. If this work has been valuable to you, consider supporting my research to help me continue building tools that benefit the entire AI community. Every contribution directly funds more open source innovation! ☕
This is the initial release of our artificial intelligence model on Hugging Face. It is important to note that this version is just the beginning; the model will be constantly improved over time. <u>Currently, the model boasts an accuracy of 79%, but we plan to increase this regularly through monthly updates.</u>
We are proud to announce that our model is currently the only one in Italy capable of identifying a wide range of 52 different categories. This capability distinctly sets it apart from other models available in the Italian landscape, offering an unprecedented level of versatility and breadth in entity recognition.
The model is based on the BERT architecture, one of the most advanced technologies in the field of Natural Language Processing (NLP). State-of-the-art techniques have been employed for its training, ensuring high-level accuracy and efficiency. This technological choice ensures a deep and sophisticated understanding of natural language.
The model is capable of identifying the following categories:
To utilize this model:
from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline
import torch
tokenizer = AutoTokenizer.from_pretrained("DeepMount00/Italian_NER_XXL")
model = AutoModelForTokenClassification.from_pretrained("DeepMount00/Italian_NER_XXL", ignore_mismatched_sizes=True)
nlp = pipeline("ner", model=model, tokenizer=tokenizer)
example = """Il commendatore Gianluigi Alberico De Laurentis-Ponti, con residenza legale in Corso Imperatrice 67, Torino, avente codice fiscale DLNGGL60B01L219P, è amministratore delegato della "De Laurentis Advanced Engineering Group S.p.A.", che si trova in Piazza Affari 32, Milano (MI); con una partita IVA di 09876543210, la società è stata recentemente incaricata di sviluppare una nuova linea di componenti aerospaziali per il progetto internazionale di esplorazione di Marte."""
ner_results = nlp(example)
print(ner_results)
The primary goal of this model is to provide effective and accurate identification of a wide range of entities, surpassing the limits of traditional models. Being the only model in Italy to recognize so many entities, we are confident that it will be an invaluable tool for numerous application areas. Constant evolution and improvement of the model is our top priority to ensure always top-notch performance.
If you are interested in contributing to this project, have suggestions for improvement, or require a specific named entity recognizer for your use case, please feel free to reach out. Your input and collaboration can significantly enhance the model's capabilities and applications. For any inquiries or to discuss potential contributions, please contact Michele Montebovi at [email protected]. Your support and participation are highly appreciated as we aim to continuously improve and expand the functionalities of the Italian_NER_XXL model.