Downloads · 30 days
15
27% of all-time downloads
THemidli/applied-ner-stage4-bert-tiny-improved
applied-ner-stage4-bert-tiny-improved is a token classification model from THemidli. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
An eight-label English token classifier fine-tuned from google/bertuncasedL-2H-128A-2. Repository: THemidli/applied-ner-stage4-bert-tiny-improved.
Downloads · 30 days
15
27% of all-time downloads
All-time downloads
55
Public
Parameters
4.4M
17.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.5 MB · 96%
From the Hugging Face model README
An eight-label English token classifier fine-tuned from google/bert_uncased_L-2_H-128_A-2. Repository: THemidli/applied-ner-stage4-bert-tiny-improved.
Exact entity-level seqeval metrics:
| Split | Precision | Recall | F1 | Token accuracy |
|---|---|---|---|---|
| Train | 0.9540 | 0.9709 | 0.9624 | 0.9948 |
| Test | 0.4261 | 0.5264 | 0.4710 | 0.8332 |
| Label | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| PERSON | 0.487 | 0.651 | 0.557 | 195 |
| ORGANIZATION | 0.216 | 0.252 | 0.233 | 147 |
| LOCATION | 0.436 | 0.545 | 0.484 | 143 |
| TIMEDATE | 0.785 | 0.832 | 0.808 | 167 |
| PRODUCT | 0.168 | 0.181 | 0.174 | 127 |
| WORKOFART | 0.136 | 0.247 | 0.176 | 97 |
| JOB | 0.664 | 0.798 | 0.725 | 99 |
| AMOUNT | 0.540 | 0.587 | 0.562 | 104 |
On 40 fresh, manually gold-labeled wild probes, exact span F1 was 0.5849 (precision 0.5439, recall 0.6327). Test F1 changed by +0.0025 versus Stage 3.
The benchmark covers tokenizer plus PyTorch CPU forward pass over 40 short probes, repeated 50 times. It is workload- and hardware-specific, not single-request latency.
PERSON, ORGANIZATION, LOCATION, TIMEDATE, PRODUCT, WORKOFART, JOB, AMOUNT using BIO encoding.
This is a 4.37M-parameter uncased two-layer BERT trained on a small, heterogeneous dataset. It is a compact baseline, not a production privacy system. Rare works/products, company-versus-product context, exact boundaries, and subword-heavy names remain weak. The 40-probe wild set is diagnostic, not a population benchmark.