Downloads · 30 days
292
80% of all-time downloads
flowxai/cee-pii
cee-pii is a token classification model from flowxai. Use it when you need labels on individual words, such as names. It is set up for gliner. The card lists the license as apache-2.0.
An open, small (~300M), multilingual span-level PII / sensitive-entity detector, weighted toward Central & Eastern European languages and financial identifiers that English-centric models miss — with first-class US/UK…
Downloads · 30 days
292
80% of all-time downloads
All-time downloads
363
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.onnx1.2 GB · 49%
From the Hugging Face model README
An open, small (~300M), multilingual span-level PII / sensitive-entity detector, weighted toward Central & Eastern European languages and financial identifiers that English-centric models miss — with first-class US/UK identifier coverage.
Given a piece of text, cee-pii returns character-offset spans, each tagged with a
type from a fixed 34-type taxonomy across three tiers (checksum-validated national
IDs, structured identifiers, and contextual entities). It is a GLiNER fine-tune, so it
extracts entities against natural-language type prompts rather than a baked-in label
head — which is exactly what lets the same weights serve two different consumers.
The product insight that shapes everything: checksum identifiers are easy mode;
contextual entities are where a privacy guard actually earns its keep. A regex can find
a CNP or an IBAN. What it cannot do is find Str. Aviatorilor nr. 12, ap. 4 next to
Ștefan Popescu next to an employer mention, without diacritics, in a noisy OCR fragment,
in six languages. That contextual span extraction is the hard part, it is where the error
mass lives, and it is the reason this model exists rather than a rulebook.
Two consumers of the same weights:
Fine-tuned from the Apache-2.0 urchade/gliner_multi-v2.1 (mDeBERTa-v3 backbone). Runs on
CPU; deployment target is consumer laptops, not a data center.
flowxai/cee-pii (this card) — GLiNER v1 fine-tune, ~300M params.flowxai/cee-pii-bench (companion dataset) — the held-out synthetic benchmark this
card is measured on. See flowxai/cee-pii-bench.This model does NOT guarantee GDPR / regulatory compliance. It is a detection aid, not a legal control, and its recall is not 100%. See Intended use & limitations before deploying.
cee-pii is a GLiNER model, so it runs through the gliner
package. You pass the text and the list of entity types you want (as natural-language
prompts); it returns spans with start / end character offsets, a label, and a score.
pip install gliner
Real Romanian example — a support-ticket line carrying a CNP, an IBAN, and a person name, typed without diacritics (the common real-world case):
from gliner import GLiNER
model = GLiNER.from_pretrained("flowxai/cee-pii")
text = (
"Buna ziua, sunt Stefan Popescu, CNP 1960714125089, "
"va rog transferati suma pe contul RO49AAAA1B31007593840000."
)
# GLiNER takes natural-language type prompts (the same phrasings the model was
# trained against). Use the short type id or a descriptive phrasing — both work.
labels = [
"person name",
"Romanian personal numeric code (CNP)",
"IBAN",
]
for ent in model.predict_entities(text, labels, threshold=0.5):
print(f"{ent['label']:40s} [{ent['start']:>3}:{ent['end']:>3}] {ent['text']}")
person name [ 16: 30] Stefan Popescu
Romanian personal numeric code (CNP) [ 36: 49] 1960714125089
IBAN [ 66: 90] RO49AAAA1B31007593840000
The (start, end) offsets are exact character ranges, so masking is a straight slice
replacement:
ents = model.predict_entities(text, labels, threshold=0.5)
masked = text
for ent in sorted(ents, key=lambda e: e["start"], reverse=True):
masked = masked[: ent["start"]] + f"[{ent['label']}]" + masked[ent["end"] :]
# -> "Buna ziua, sunt [person name], CNP [Romanian personal numeric code (CNP)], ..."
Notes for callers.
cnp → "Romanian personal numeric code (CNP)"), so
the model responds to descriptive prompts. At evaluation we use one fixed canonical
phrasing per type and map it back to the short id — the full canonical list is the
taxonomy below.0.5 is the evaluated operating point (0.911
exact micro-precision, 0.758 recall on the bench). Lower it for a higher-recall guard;
raise it for a stricter masking layer.max_types was 25 at training time; pass entity lists in batches if you need all 34.text --> cee-pii (GLiNER, mDeBERTa-v3 encoder + span head) --> [ (type, start, end, score), ... ]
^
| entity-type prompts (natural-language phrasings, open-vocabulary)
Under the hood: whitespace word-splitter → mDeBERTa-v3 encoder → span scorer against each type prompt → threshold → char-offset spans. GLiNER scores candidate word-token spans
(max_width 12) against each supplied type prompt; spans above threshold are emitted
with their character offsets recovered from the splitter. There is no fixed classification
head, which is why the taxonomy can grow without retraining the output layer.
The canonical type list is the single source of truth shared by the generators, validators, corpus, bench, and eval adapter — it cannot drift. Grouped by the three design tiers.
| Type | Country | Notes |
|---|---|---|
cnp | RO | 13-digit personal numeric code; control digit; encodes DOB + county |
ci_ro | RO | ID card series (2 letters, valid county) + 6 digits |
pesel | PL | 11 digits, positional checksum, encodes DOB |
nip | PL | 10-digit tax ID, weighted checksum |
taj | HU | 9-digit social-security number, checksum on first 8 |
szemelyi | HU | 11-digit personal ID (személyi szám) |
pinfl | UZ | 14-digit personal ID (official lex.uz checksum) |
iban | RO/PL/HU/GB + generic | ISO 13616 mod-97, correct country lengths |
card | intl | Luhn + realistic scheme shapes |
nhs | UK | 10 digits, mod-11 checksum |
aba | US | 9-digit bank routing number, 3-7-1 weighted checksum |
| Type | Country | Notes |
|---|---|---|
ssn | US | 9 digits; invalid-range rejection (area 000/666/900+, group 00, serial 0000) |
itin | US | 9xx- with valid IRS group ranges |
ein | US | valid prefix list |
nino | UK | National Insurance number (prefix + suffix rules) |
utr | UK | 10-digit Unique Taxpayer Reference (HMRC check digit) |
company_number_uk | UK | 8-char Companies House number (incl. SC/NI) |
uk_sort_code | UK | xx-xx-xx |
uk_account_number | UK | 8 digits |
uz_account | UZ | 20-digit domestic account number (structural) |
phone | multi | E.164 + local formats per country |
email | — | email address |
plate | RO/PL/HU/UK | vehicle registration |
postal | multi | postal / ZIP (incl. UK postcode grammar, US ZIP+4) |
dob | multi | date of birth, many formats |
| Type | Notes |
|---|---|
person_name | full name (with AND without diacritics) |
first_name | given name only |
surname | family name only |
address | street address (incl. RO Str./nr./bl./sc./et./ap. shape) |
policy_ref | insurance policy number / reference |
contract_ref | contract number / reference |
account_ref | internal account number / reference (not a bank IBAN) |
employer | employer / company mention |
health_condition | coarse flag that a health condition is mentioned — not classified |
All numbers below are on the synthetic CEE-PII-Bench v0.2 — read the honesty caveat first, then the tables.
Honesty caveat — home-field advantage (read before citing these numbers). CEE-PII-Bench v0.2 is held-out and contamination-verified (bench families ∩ train = ∅ and bench values ∩ train = ∅, both tested), but it is drawn from the same synthetic generator distribution as the training corpus — the same output format, phrasing, and entity style the fine-tune learned. The numbers here are a valid relative signal (fine-tune vs zero-shot vs frontier on identical inputs) but are likely optimistic in absolute terms versus real documents. This is the same caveat as the sibling scam-guard project. The honest real-world number is the Phase-6 human-labeled real-structure eval set (official form specimens, published template contracts, sample bank statements), which is reported separately and is still to be built.
Scoring is model-agnostic (eval/harness.py): each entity is a (type, start, end)
triple, matched greedily one-to-one per type. Exact requires the span boundaries and
type to match; relaxed requires the type to match and the character ranges to overlap.
The frontier column runs the same task through the same scorer (LLM offsets recovered
by verbatim-substring search, since raw LLM char offsets are unreliable).
Full 2,000-doc v0.2 standard bench (reports/eval_gliner.md). The Claude column is on
a 100-doc seeded slice (the cached frontier reference, reports/eval_frontier.md); the
fine-tuned model's score on that same 100-doc slice is shown in the Claude row's note
for a fair comparison.
| Metric | zero-shot gliner_multi-v2.1 | fine-tuned cee-pii v1 | lift vs zero-shot |
|---|---|---|---|
| micro-F1 (exact) | 0.177 | 0.827 | +0.650 (4.7×) |
| micro-F1 (relaxed) | 0.246 | 0.833 | +0.587 (3.4×) |
| macro-F1 (exact) | 0.114 | 0.598 | +0.484 (5.2×) |
| micro-precision (exact) | 0.661 | 0.911 | +0.250 |
| micro-recall (exact) | 0.102 | 0.758 | +0.656 |
| FP-rate (274 no-entity docs) | 0.000 | 0.000 | matched |
The headline is the false-positive rate: 0.000 on the 274 no-entity documents. A privacy guard that fires on clean text gets disabled within a week; this one does not fire on documents that carry no PII, while still reaching 0.758 recall at 0.911 precision.
Acceptance gate (Phase 4): MET — the fine-tune beats zero-shot GLiNER by a wide margin (+0.65 exact micro-F1, a 4.7× relative gain). Zero-shot's failure is almost entirely recall (0.10): the base model fires on only a handful of universal types (postal, dob) and ignores the CEE-specific taxonomy. Fine-tuning is what teaches the taxonomy.
Gap to the Claude frontier (reference, not a competitor). On the same 100-doc slice as
the cached Claude Opus 4.8 column, fine-tuned cee-pii scores 0.833 / 0.841
(exact / relaxed micro-F1) vs Claude's 0.936 / 0.963 — a ~0.10 exact-F1 gap
(reports/eval_gliner_slice100.md). Expected and honest: a ~300M open-weights,
CPU-deployable, Apache-2.0 model reaching ~89% of a frontier API's exact-F1 on this
bench, fully offline and at a fraction of the cost and latency. Not an apples-to-apples
comparison — the value proposition is on-device masking, not beating a frontier LLM.
| Language | exact micro-F1 |
|---|---|
| en_uk | 0.95 |
| pl | 0.94 |
| en_us | 0.89 |
| uz | 0.74 |
| ro | 0.66 |
| hu | 0.63 |
RO and HU trail — and honestly so. Their weakest types (ci_ro, taj, szemelyi)
concentrate in those languages, so the per-type errors below drag the per-language number
down. Closing the RO/HU gap is the explicit target of the planned 3rd-epoch follow-up.
cnp 1.00, phone/email/postal ~0.99,
person_name 0.98, dob 0.98, card 0.97, employer/plate 0.96, utr 0.94.ci_ro (RO ID card) F1 0.00 — the model finds the span but mislabels it as
nino (UK NI number); both are "2 letters + digits", a learnable confusion.uz_account 0.00 — a genuine recall miss on a rare, under-represented
long-digit type.taj 0.17 — confused with generic numeric references (policy_ref).first_name 0.57 / surname 0.62 — boundary/role confusion with
person_name.ein, itin, nino, pesel, ssn,
szemelyi, uk_account_number show F1 0.00 but have zero gold occurrences in v0.2
standard — the 0.00 is a scoring convention (P=R=0 when no gold exists), not a model
failure. v0.2 standard exercises 23 of the 34 taxonomy types; the remaining types
need bench coverage before their F1 is meaningful.The full per-type / per-language tables (fine-tuned + zero-shot + heuristic floor) live in
reports/eval_gliner.md; the 100-doc frontier slice is in
reports/eval_gliner_slice100.md and
reports/eval_frontier.md.
Hard-subset (960-doc noised) eval, XLM-R BIO baseline + Presidio through the same scorer, CPU latency + peak memory on M3-class hardware, bootstrap confidence intervals over documents, the Phase-6 human-labeled real-structure eval set, and OpenAI/Gemini frontier columns once keys are configured.
Intended use. A detection aid for (1) redacting PII before text leaves a perimeter or reaches an LLM, and (2) warning a user before they paste PII into a chatbot. Languages: Romanian, Polish, Hungarian, Uzbek, UK English, US English. Deployment target is consumer CPU.
Out of scope & limitations.
ci_ro↔nino, taj↔policy_ref, szemelyi) concentrate there.All training data is synthetic. Zero client data, zero scraped personal data, zero real PII anywhere in the repo (including tests).
claude-haiku-4-5, slot placeholders preserved
and validated, all outputs cached for reproducibility).cee-pii-phase3-v0.2, seed 20260702),
balanced per language (each language 16.3–17.2%), ~14% zero-entity docs, hard-negative
injection in ~50% of docs, and noise applied at assembly (diacritic stripping — critical
for RO — OCR confusions, random casing, whitespace/punctuation damage) with
character-level span tracking through every transformation.(language, register), seed
20260703, 80/10/10), then each split is generated only from its own family pool →
straddler_count = 0 by construction. Value-disjointness enforced by resample. Split
sizes: train 16,000 / val 2,000 / test 2,000.Full fine-tune of urchade/gliner_multi-v2.1 (mDeBERTa-v3 backbone, ~300M params) on
corpus v0.2 train (config training/config/gliner_v1_2ep_memsafe.yaml, run
training/runs/gliner_v1/).
cnp → "Romanian personal numeric code (CNP)"), 2–3 phrasings per
type sampled per document to preserve the open-vocabulary property.max_width 12, max_types 25; seed 20260704. Effective
batch 2 × 16 = 32 (small physical batch is a deliberate MPS memory-safety choice,
not a quality preference).training/runs/gliner_v1/final (end of epoch 2).
PYTORCH_ENABLE_MPS_FALLBACK=1 set so any unsupported op degrades to CPU rather than
crashing a multi-hour run.Base-model license.
urchade/gliner_multi-v2.1is Apache-2.0 (verified on its HF model card, 2026-07-05), compatible with this Apache-2.0 release.
The model repo carries the promoted GLiNER checkpoint from training/runs/gliner_v1/final:
pytorch_model.bin — the fine-tuned weights (~1.15 GB).gliner_config.json — GLiNER config (mDeBERTa-v3 encoder, max_width 12, span mode
markerV0, etc.).tokenizer.json + tokenizer_config.json — the mDeBERTa-v3 tokenizer.GLiNER.from_pretrained("flowxai/cee-pii") loads these directly.
onnx/model.fp32.onnx, an fp32 ONNX export, is also published. GLiNER's own
export_to_onnx() traces the model's LSTM span head through pack_padded_sequence /
pad_packed_sequence, which does not survive tracing and produces an invalid graph.
The export here works around that with a patched, unpacked LSTM forward pass that is
provably equivalent to the packed one for a single, unpadded sequence, which is the
only shape this model is ever called with in practice: batch size 1, one full-length
text, no padding. Verified against the unmodified PyTorch model on 7 fixtures across
en/ro/pl/hu, 0 span mismatches, maximum score drift 0.00001. onnx/export_manifest.json
carries the fixtures and the numbers. The export and verification script,
gliner_to_onnx.py, is the source this repeats.
No int8 export is offered. Naive dynamic INT8 quantization of this model destroys
it: every real entity name's score dropped below 0.004 on both the qnnpack and
fbgemm backends. fp16 was also attempted and not shipped, for a narrower reason: one
conversion bug turned out to be a metadata mismatch with no effect on computation, but
a second is a genuine operand dtype mismatch in mDeBERTa's embeddings block that the
converter does not resolve. fp32 is what ships until a quantisation recipe is found
that this architecture tolerates.
# environment (uv-managed, Python 3.12)
uv sync
# corpus v0.2 (split-aware generation, family- + value-disjoint, seed 20260702)
uv run python scripts/build_corpus_v02.py
# freeze CEE-PII-Bench v0.2 from the test split (refuses silent overwrite)
uv run python scripts/build_bench.py
# training: GLiNER fine-tune (Phase 4, ~1.5h on M3 Max MPS)
uv run python training/train_gliner.py --config training/config/gliner_v1_2ep_memsafe.yaml
# eval: fine-tuned vs zero-shot GLiNER on the full 2,000-doc bench
uv run python scripts/run_eval_gliner.py --full # -> reports/eval_gliner.md
# frontier reference on a 100-doc slice (Claude; openai/gemini skip cleanly w/o keys)
env -u ANTHROPIC_BASE_URL -u ANTHROPIC_AUTH_TOKEN \
uv run python scripts/run_eval_frontier.py --n 100 --seed 20260703 \
--specs claude:claude-opus-4-8
flowxai/cee-pii-bench.Apache-2.0 (weights, code, and data pipeline). Base model urchade/gliner_multi-v2.1 is
Apache-2.0.