Downloads · 30 days
336
61% of all-time downloads
patronus-studio/orca-sonar-document-classifier
orca-sonar-document-classifier is a text classification model from patronus-studio. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
Multilingual Document Topic Classifier for Real-World AI Security & DLP
Downloads · 30 days
336
61% of all-time downloads
All-time downloads
554
Public
Parameters
141M
3.3 GB on disk
Likes
4
Public
Click a slice to open those files.
.onnx1.3 GB · 59%
From the Hugging Face model README
Multilingual Document Topic Classifier for Real-World AI Security & DLP
Read more
Orca-Sonar is a Multilingual ModernBERT-based (mmBERT) classifier that assigns a document/text to one of 9 topic classes. It is part of the Patronus Protect security stack and is designed for topic-/risk-routing of incoming texts (e.g. before they reach an LLM, a DLP gate, or a storage tier).
It classifies German and English text and is robust to user-to-AI wrappers (e.g. "Summarize this contract: …"), i.e. the topic of the content determines the class, not the surface format of the request.
The model maps an input text to one of:
| id | label | description |
|---|---|---|
| 0 | legal | contracts, NDAs, ToS/AGB, privacy policies, statutes/judgments, compliance, legal correspondence |
| 1 | hr | CVs, job ads, employment contracts, terminations, HR policies, performance reviews, recruiting |
| 2 | finance | invoices, balance sheets, quarterly/annual reports, cash-flow, SEC filings, forecasts |
| 3 | internal_and_tech | ADRs, RFCs, postmortems, specs, READMEs, wikis, architecture & strategy memos, runbooks |
| 4 | source_code | raw program code & configs (Python/Go/Rust/JS/TS/SQL/Bash/Dockerfile/k8s/Terraform …) |
| 5 | marketing | press releases, newsletters, landing-page/sales copy, outbound pitches, case studies |
| 6 | other | conversational / non-business: smalltalk, recipes, travel, hobby, learning, creative |
| 7 | education | curricula, course materials, transcripts, grades, assignments, and academic records |
| 8 | medical | clinical notes, diagnoses, treatments, medications, and healthcare records |
Disambiguation: on a tie, the more sensitive class wins:
legal > hr > finance > internal_and_tech > source_code > marketing > other.
model.safetensors, fp32).onnx/onnx_fp16/, half the size, argmax-faithful to the full model.Trained on our own in-house dataset (German + English, 9 topic classes), purpose-built for this model. The dataset will be published soon.
Held-out test set (100 % real data), per-class F1:
| Metric | Score |
|---|---|
| Accuracy | 0.942 |
| F1 (macro) | 0.940 |
| F1 medical | 0.961 |
| F1 education | 0.957 |
| F1 marketing | 0.956 |
| F1 finance | 0.948 |
| F1 source_code | 0.945 |
| F1 other | 0.943 |
| F1 legal | 0.942 |
| F1 internal_and_tech | 0.930 |
| F1 hr | 0.877 |
from transformers import pipeline
clf = pipeline("text-classification", model="patronus-studio/orca-sonar-document-classifier")
clf("Fasse mir diesen Dienstleistungsvertrag zusammen: Laufzeit 24 Monate, Gerichtsstand München …")
# -> [{'label': 'legal', 'score': 0.99}]
An FP16 ONNX version is available under onnx/onnx_fp16/:
import torch
from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer
model_id = "patronus-studio/orca-sonar-document-classifier"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForSequenceClassification.from_pretrained(model_id, subfolder="onnx/onnx_fp16")
inputs = tokenizer("def add(a, b):\n return a + b", return_tensors="pt")
logits = model(**inputs).logits
print(model.config.id2label[int(torch.argmax(logits, dim=-1))])
@misc{orcasonar2026,
title={Orca-Sonar: Multilingual Document Topic Classification for Real-World AI Security},
author={Patronus Protect},
year={2026},
howpublished={\url{https://huggingface.co/patronus-studio/orca-sonar-document-classifier}}
}
This model is released under the Apache License 2.0.
A copy of the license is included as LICENSE in this repository.
The model is derived from jhu-clsp/mmBERT-small, which is distributed under the MIT License. The upstream copyright and permission notice are retained; the MIT terms continue to apply to the portions originating from that work.
This model is built to run inside Patronus Ark, Patronus' open-source on-device AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer). Ark is open source: GitHub repository · product page.
Brought to you by Patronus Protect, a local AI firewall that secures every AI interaction (prompts, tools, documents) before it reaches your models. Try it for free at patronus.studio.