Downloads · 30 days
16
25% of all-time downloads
MatteoFasulo/ModernBERT-large-NER
ModernBERT-large-NER is a token classification model from MatteoFasulo. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of answerdotai/ModernBERT-large for Named Entity Recognition (NER) tasks on conll2003 dataset.
Downloads · 30 days
16
25% of all-time downloads
All-time downloads
65
Public
Parameters
396M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of answerdotai/ModernBERT-large for Named Entity Recognition (NER) tasks on conll2003 dataset.
ModernBERT-large-NER is a token classification model trained to identify and categorize named entities in text. Built on the ModernBERT-large architecture, this model leverages modern transformer optimizations for efficient and accurate entity extraction.
Primary Use Cases:
Intended Users:
Known Limitations:
Out-of-Scope Uses:
The model was trained on a dataset for named entity recognition. Specific details about the dataset composition, size, and entity types are not publicly disclosed in this release.
It achieves the following results on the evaluation set:
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
|---|---|---|---|---|---|---|---|
| No log | 1.0 | 439 | 0.0776 | 0.8749 | 0.9122 | 0.8931 | 0.9800 |
| 0.1518 | 2.0 | 878 | 0.0508 | 0.9230 | 0.9399 | 0.9314 | 0.9861 |
| 0.0334 | 3.0 | 1317 | 0.0509 | 0.9219 | 0.9493 | 0.9354 | 0.9880 |
| 0.0097 | 4.0 | 1756 | 0.0535 | 0.9267 | 0.9505 | 0.9384 | 0.9888 |
| 0.0029 | 5.0 | 2195 | 0.0555 | 0.9272 | 0.9519 | 0.9394 | 0.9889 |
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer, pipeline
# Create NER pipeline
ner_pipeline = pipeline(
"token-classification",
model="MatteoFasulo/ModernBERT-large-NER",
aggregation_strategy="simple",
dtype=torch.bfloat16,
)
# Example usage
text = "Apple Inc. was founded by Steve Jobs in Cupertino, California."
entities = ner_pipeline(text)
for entity in entities:
print(
f"{entity['word']}: {entity['entity_group']} (confidence: {entity['score']:.4f})"
)
# Apple Inc.: ORG (confidence: 0.9684)
# Steve Jobs: PER (confidence: 0.9950)
# Cupertino: LOC (confidence: 0.9876)
# California: LOC (confidence: 0.9939)
Privacy: This model may extract personal information (names, locations, organizations) from text. Users should:
Bias: The model's performance may reflect biases present in the training data, potentially affecting:
Users should validate the model's performance on their specific use cases and implement bias mitigation strategies as needed.
If you use this model in your research, please cite ModernBERT model:
@misc{modernbert,
title={Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
author={Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
year={2024},
eprint={2412.13663},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.13663},
}
This model is released under the Apache 2.0 License. See the LICENSE file for details.
This model was built using the ModernBERT-large architecture from Answer.AI and trained using the Hugging Face Transformers library.