Downloads · 30 days
0
pidakwo/rtc-ner-extended
rtc-ner-extended is a token classification model from pidakwo. Use it when you need labels on individual words, such as names. It is set up for spacy.
rtc-ner-extended is a domain-specific Named Entity Recognition (NER) model developed for extracting structured information from unstructured road traffic crash (RTC) narratives. It works to identify entities not cover…
Downloads · 30 days
0
Access
Public
Updated Sep 2, 2026
Repo size
6.1 MB
Likes
0
Public
Click a slice to open those files.
Other6.2 MB · 98%
From the Hugging Face model README
rtc-ner-extended is a domain-specific Named Entity Recognition (NER) model developed for extracting structured information from
unstructured road traffic crash (RTC) narratives. It works to identify entities not covered by the pidakwo/rtc-ner model
The model identifies RTC attributes such as rtc_time, weekday, no_vehicle, vehicle_type, persons_involved, and rtc_factors (causal
factors, collision types, collision effects, physical site attributes, environmental condition, actors – humans and animals, goods types,
and general vehicle category). Additionally, the model was also used as part of a data anonymization workflow in which it was used to identify
entities - person, plate_no, address, and victim_organization which indicate personal details of road users involved in the RTC incident.
In order to hide the personal details in the content feature, NER entities identified were replaced with their label names.
rtc-ner-extended is an English-language spaCy NER model consisting of a tok2vec component and an NER component.
The model uses spaCy's:
MultiHashEmbed representation for token featuresMaxoutWindowEncoder for contextual token representationsThe model was configured and trained using spaCy 3.8.x.
rtc-ner-extended recognizes the following ten domain-specific entity types:
| Entity | Description |
|---|---|
ADDRESS | Address of RTC incident victims appearing in an RTC narrative |
FACTORS | Reported factors or circumstances associated with the occurrence of the crash |
NO_VEHICLES | Number of vehicles involved in the crash |
PERSON | Name of a person mentioned in the RTC narrative |
PERSONS_INVOLVED | Total number of persons involved in the crash |
PLATE_NO | Vehicle registration or license plate number |
TIME | Time information associated with the crash or reported event |
VEHICLE_TYPE | Type or category of vehicle involved in the crash |
VICTIM_ORGANIZATION | Organization associated with a victim or incident |
WEEKDAY | Day of the week associated with the crash |
rtc-ner-extended is intended primarily for research and information-extraction applications involving road traffic crash narratives.
Potential applications include:
rtc-ner-extended was also used as part of a privacy-preserving data processing workflow.
Four entity categories were specifically identified as potentially containing personally identifiable information (PII):
PERSONPLATE_NOADDRESSVICTIM_ORGANIZATIONThese entities can be identified in an RTC narrative and subsequently replaced with their entity labels or other designated placeholders during anonymization.
Important: NER-based anonymization should not be regarded as a guarantee that all personally identifiable or sensitive information has been removed. Human review and/or additional privacy-preserving processing should be considered before publicly releasing processed text.
The model uses the following spaCy pipeline:
tok2vec → ner
The repository contains the complete trained spaCy model and its associated resources, including:
config.cfg
meta.json
tokenizer
vocab/
ner/
tok2vec/
These components should be retained together when loading the model.
The performance figures are recorded in the model's meta.json metadata.
The RTC_NER_Extended_model_test.ipynb file contains the code for testing usage of the model.
The final output contains the entities identified by the model together with their corresponding entity labels.
rtc-ner-extended is a domain-specific research model and should not be assumed to identify every relevant entity in every RTC narrative.
Performance may be affected by:
In particular, automated anonymization should not be considered sufficient on its own to guarantee that an RTC narrative contains no personally identifiable information.
For applications involving public release of textual data, model predictions should be complemented by appropriate validation and privacy review.
rtc-ner-extended was developed as part of research investigating the transformation of unstructured road traffic crash narratives into structured, machine-readable information.
The model extends the information-extraction capability of the domain-specific pidakwo/rtc-ner model pipeline by identifying entities associated with crash
circumstances, persons, vehicles, temporal information, and potentially privacy-sensitive information.
The extracted information can subsequently support data curation, analysis, machine learning, and road safety research.
A separate model, pidakwo/rtc-ner, was developed for the extraction of geographic and selected incident-related entities from RTC narratives.
RTC-NER and rtc-ner-extended are separate but complementary models, and should not be treated as interchangeable.
The code for model training and evaluation as well as data extraction can be found at: https://github.com/PatUnoka/Geospatial-and-Contextual-Information-Extraction-from-Road-Traffic-Crash-Narratives.git
If you use rtc-ner in academic research, please cite the associated research publication and dataset from which the model was developed.
[1] P. O. Idakwo, O. Adekanmbi, A. Soronnadi, and A. David, “Geo-parsing and analysis of road traffic crash incidents for data-driven emergency response planning,” Heliyon, vol. 11, no. 4, p. e41067, 2025, doi: 10.1016/j.heliyon.2024.e41067.
[2] P. O. Idakwo, O. Adekanmbi, and A. David, “Nigerian Multi-modal Road Traffic Crash Data,” 2026, doi: 10.5281/ZENODO.15862127.