Downloads · 30 days
7
12% of all-time downloads
sammmmmmm25/alfred-ner-bert
alfred-ner-bert is a token classification model from sammmmmmm25. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
This is the entity-extraction subcomponent of Alfred AI Investigations, a Graph RAG assistant for reasoning across criminal case documents. It is not intended to be used standalone -- see the main project repository l…
Downloads · 30 days
7
12% of all-time downloads
All-time downloads
57
Public
Parameters
108M
431 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors431 MB · 100%
From the Hugging Face model README
This is the entity-extraction subcomponent of Alfred AI Investigations, a Graph RAG assistant for reasoning across criminal case documents. It is not intended to be used standalone -- see the main project repository linked above for full context, architecture, and evaluation results.
A fine-tuned version of dslim/bert-base-NER, retrained to recognize 8 entity types relevant to investigative case documents rather than the base model's original 4 general-purpose types:
| Label | Description |
|---|---|
PER | Person names |
ORG | Organizations |
LOC | Locations |
DATE | Dates and times |
VEH | Vehicle descriptions |
WEAPON | Weapons |
CASE | Case numbers |
ALIAS | Aliases / nicknames / saved contact names |
At inference time, extracted entities are used to build a chunk-level co-occurrence graph: document chunks sharing an entity (by exact match or fuzzy word-overlap match for multi-word spans) are linked, enabling retrieval to traverse cross-document connections that similarity search alone would miss.
Fine-tuned on a small hand-authored, BIO-tagged dataset (~16-20 examples) covering all 8 entity types, written to match the vocabulary of Alfred's synthetic case corpus (witness names, vehicle descriptions, call log entries, case numbers).
dslim/bert-base-NER| Metric | Value |
|---|---|
| Precision | 0.778 |
| Recall | 0.778 |
| F1 | 0.778 |
| Accuracy | 0.923 |
These numbers are modest and expected given the small training set -- see Limitations below and the main project's Limitations section for the practical consequences this has on downstream retrieval quality.
Component of the Alfred AI Investigations pipeline only. Not intended as a general-purpose NER model, and not validated for use in real investigations without substantial additional training data and evaluation.