Downloads · 30 days
0
SanoAI/sano-shield-1-multilingual
sano-shield-1-multilingual is a token classification model from SanoAI. Use it when you need labels on individual words, such as names. It is set up for transformers.
Sano Shield 1.2 is Sano AI's local privacy model for finding sensitive spans in German and English before text reaches an external AI provider. Its stable technical model ID is sano-shield-1-multilingual.
Downloads · 30 days
0
Access
Public
Updated Aug 23, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md11.7 KB · 88%
From the Hugging Face model README
Sano Shield 1.2 is Sano AI's local privacy model for finding sensitive spans in German
and English before text reaches an external AI provider. Its stable technical model ID is
sano-shield-1-multilingual.
Shield 1.2 was optimized for a deliberately narrow product goal: reduce false positives without giving back required protection. Compared with Shield 1.1 on the same frozen German and English product suites, unknown false positives fall by 17.0% while more required spans are protected in both languages.
Availability: This Hugging Face repository publishes the model card for transparency. Model weights, tokenizer files, ONNX artifacts, and training data are not publicly distributed or downloadable from this repository.
Evaluation scope: All reported quality suites are internal and synthetic. They were held out from training, but they are not a substitute for independently sourced, human-reviewed, domain-specific evaluation. These results do not establish a global state-of-the-art claim.
| Field | Value |
|---|---|
| Model family | Sano Shield |
| Release | Sano Shield 1.2 |
| Technical ID | sano-shield-1-multilingual |
| Task | Token classification / sensitive-entity recognition |
| Evaluated languages | German and English |
| Entity classes | 35 entity types, 69 BIO token labels |
| Architecture | XLM-RoBERTa token classifier |
| Base-model lineage | bardsai/eu-pii-anonimization-multilang |
| Pinned base revision | 0e72e19f030ed4e661b1673e549af8e0dd176386 |
| Release date | 2026-08-23 |
| Deployment format | ONNX with FP16 word embeddings and FP32 transformer/classifier |
| Deployment size | 726,390,468 bytes |
| Deployment ONNX SHA-256 | 7bb99023a369be99cd5e4bf949a2d8bc2999bb6e9c67569dd58fe52f4adbe86e |
Shield 1.2 starts from the private Shield 1.1 checkpoint and applies targeted continuation training on Sano Shield v0.7. The selected release is a weight interpolation:
Shield 1.2 = 0.25 × Shield 1.1 + 0.75 × targeted continuation
The continuation used seed 20260824. Embeddings and the first four encoder layers were
frozen, and a Shield 1.1 teacher supplied a small preservation loss. Candidate selection
required separate German and English quality gates, privacy-span preservation, hard-negative
behavior, ONNX parity, the current desktop product pipeline, and document usability.
The strongest raw continuation reduced more synthetic false positives, but was not selected for deployment because it failed frozen organization and per-type product preservation. The interpolated release is the strongest tested artifact that passes every release gate.
Shield 1.2 is designed as one component in a reversible privacy pipeline:
Production deployments should retain deterministic recognizers, overlap resolution, user-visible review, monitoring, and fail-closed handling. Organization names remain user-selectable but are reported separately from mandatory natural-person PII in the product release policy.
Do not use the model as the sole basis for legal-compliance decisions, medical decisions, identity verification, employee monitoring, or irreversible actions affecting a person.
The frozen Sano Shield v0.7 test split contains 5,000 German and English records. This benchmark isolates the neural model; it does not include deterministic recognizers or the desktop's overlap-resolution logic.
| Scope | Model | Entity F1 | Precision | Recall | False positives | False negatives |
|---|---|---|---|---|---|---|
| Combined | Shield 1.1 | 0.7962 | 72.55% | 88.21% | 4,517 | 1,596 |
| Combined | Shield 1.2 | 0.8286 | 75.78% | 91.41% | 3,955 | 1,162 |
| German | Shield 1.1 | 0.7837 | 70.50% | 88.22% | 2,477 | 791 |
| German | Shield 1.2 | 0.8165 | 73.79% | 91.39% | 2,179 | 578 |
| English | Shield 1.1 | 0.8088 | 74.68% | 88.20% | 2,040 | 805 |
| English | Shield 1.2 | 0.8409 | 77.84% | 91.44% | 1,776 | 584 |
Overall, Shield 1.2 produces 562 fewer false positives and 434 fewer false negatives. The largest targeted F1 gains are AUTH_SECRET (+0.271), PERSON_ALIAS (+0.213), CONTACT_HANDLE (+0.183), ACCOUNT_IDENTIFIER (+0.130), and DOCUMENT_REFERENCE (+0.116).
The model-only benchmark also exposes regressions that remain important even though the complete product pipeline passes:
| Watch metric | Shield 1.1 | Shield 1.2 | Change |
|---|---|---|---|
| Privacy-span recall | 99.889% | 99.793% | -0.096 percentage points |
| Leak-free record rate | 99.624% | 99.298% | -0.326 percentage points |
| LOCATION F1 | 0.6915 | 0.6781 | -0.0134 |
| PERSON_ROLE_OR_TITLE F1 | 0.9719 | 0.9548 | -0.0171 |
The release-relevant benchmark runs the exact deployment artifact through Sano's current judge-free desktop pipeline. It uses paired held-out synthetic stress suites with 3,000 records per language and includes deterministic recognizers and overlap resolution.
| Language | Model | Required spans protected | Required-span recall | Complete positive records | Records without a required miss | Unknown false positives | Optional organizations |
|---|---|---|---|---|---|---|---|
| German | Shield 1.1 | 5,774/5,896 | 97.93% | 2,150/2,272 | 2,878/3,000 | 218 | 1,540/1,540 |
| German | Shield 1.2 | 5,796/5,896 | 98.30% | 2,172/2,272 | 2,900/3,000 | 182 | 1,540/1,540 |
| English | Shield 1.1 | 5,903/6,149 | 96.00% | 2,034/2,280 | 2,754/3,000 | 363 | 1,590/1,590 |
| English | Shield 1.2 | 5,923/6,149 | 96.32% | 2,054/2,280 | 2,774/3,000 | 300 | 1,590/1,590 |
| Combined | Shield 1.1 | 11,677/12,045 | 96.94% | 4,184/4,552 | 5,632/6,000 | 581 | 3,130/3,130 |
| Combined | Shield 1.2 | 11,719/12,045 | 97.29% | 4,226/4,552 | 5,674/6,000 | 482 | 3,130/3,130 |
The product result is the basis for promotion: 42 more required spans are protected and unknown false positives fall by 99 (-17.0%). Every frozen German, English, mandatory-type, organization-preservation, Tier-A, and aggregate false-positive gate passes.
The current no-judge benchmark contains 30 mixed-format documents covering TXT, Markdown, CSV, JSON, DOCX, XLSX, PPTX, text-layer PDF, and OCR-derived PDF text.
| Measure | Shield 1.1 | Shield 1.2 |
|---|---|---|
| Zero-touch documents | 17/30 | 18/30 |
| Review effort points | 17 | 15 |
| Additions | 0 | 0 |
| Tier-A additions | 0 | 0 |
| Office structures preserved | 3/3 | 3/3 |
| PDF page structures preserved | 6/6 | 6/6 |
The historical 29/30 Shield 1.1 figure used a now-retired plausibility judge and is not a valid baseline for the current architecture. Both versions above were replayed through the same current pipeline.
Runtime figures are medians of three fresh processes per model on the same 14-core macOS host with ONNX Runtime 1.27.0 and the CPU execution provider. Build time is excluded.
| Measure | Shield 1.1 | Shield 1.2 | Change |
|---|---|---|---|
| Cold model startup | 425.1 ms | 444.6 ms | +4.6% |
| Short-window mean | 16.61 ms | 16.32 ms | -1.7% |
| Long-window mean | 141.40 ms | 143.57 ms | +1.5% |
| Startup RSS delta | 1,511.6 MB | 1,511.4 MB | -0.01% |
| 32-page detection | 12.414 s | 12.513 s | +0.8% |
| 32-page cold end-to-end | 18.811 s | 19.053 s | +1.3% |
| Protected occurrences in fixture | 808 | 812 | +4 |
All six protected-copy runs completed with placeholders and zero selected-original leaks. The PDF fixture is a performance and smoke-test corpus, not a labeled accuracy corpus.
The FP16-embedding ONNX artifact was compared with its FP32 reference on 512 held-out v0.7 test records (23,662 active tokens):
202608247.5e-7 learning rate, 5% warmup[0.5, 2.5], outside-label
weight 1.25Training targets come from Sano's versioned synthetic corpus, not predictions copied from a separate evaluator. Customer content and real personal data are not accepted into the training repository.
This public page documents the model's purpose, lineage, evaluation, and limitations. It does not grant access to the checkpoint or corpus, and no public redistribution license for those private artifacts is granted here.