Downloads · 30 days
14
18% of all-time downloads
nielsaxe/BookTitleNERDutch
BookTitleNERDutch is a token classification model from nielsaxe. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This Named Entity Recognition (NER) model is designed to extract book titles from Dutch texts.
Downloads · 30 days
14
18% of all-time downloads
All-time downloads
78
Public
Parameters
559M
2.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
This Named Entity Recognition (NER) model is designed to extract book titles from Dutch texts.
The model has been fine-tuned and evaluated on a Dutch dataset consisting of 12,535 book reviews from the Leeuwarder Courant, identifying 23,529 book titles. The dataset utilizes the IO Tagging Schema. The data was divided into a training set (70%), validation set (15%), and test set (15%). Training involved the Majority or Minority loss function, achieving an F1 score of 84.3%, Precision of 83.4%, and Recall of 85.2% on the test set.

This model is intended for extracting book titles from Dutch texts, particularly useful for applications involving text analysis in the literary domain.
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
# Load the model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("nielsaxe/BookTitleNERDutch")
model = AutoModelForTokenClassification.from_pretrained("nielsaxe/BookTitleNERDutch")
# Create a NER pipeline
nlp = pipeline("ner", model=model, tokenizer=tokenizer)
# Example usage
text = "Gisteren heb ik het boek Nijntje in de dierentuin gelezen. Ik kan niet anders zeggen dat dit boek fantastisch was!"
entities = nlp(text)
print(entities)