Downloads · 30 days
19
3% of all-time downloads
AmirMohseni/modernbert-primary_topic
modernbert-primary_topic is a text classification model from AmirMohseni. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Fine-tuned ModernBERT-base classifier that assigns a primary legal topic to a multi-turn conversation.
Downloads · 30 days
19
3% of all-time downloads
All-time downloads
661
Public
Parameters
150M
28.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors598 MB · 99%
From the Hugging Face model README
Fine-tuned ModernBERT-base classifier that assigns a primary legal topic to a multi-turn conversation.
Part of the Legal QA collection · Try the interactive demo →
Stage 2 of a two-model encoder routing pipeline:
| Stage | Model | Input | Output |
|---|---|---|---|
| 1 | modernbert-seeks_guidance | Full conversation | seeks_legal_guidance (True/False) |
| 2 | modernbert-primary_topic | Full conversation | Primary topic (14 labels + non-guidance) |
modernbert-primary_topic predicts one of 14 legal topic labels, plus a (non-guidance) class for conversations where the user is not seeking legal help. In practice, run modernbert-seeks_guidance first and only trust the topic label when it predicts True.
Input preprocessing: full user–assistant conversation, serialized as Role: content lines per turn.
| Topic | Description |
|---|---|
FAMILY | Marriage, divorce, child custody, child support, alimony, adoption, guardianship, domestic violence, parentage, family-status disputes. |
HOUSING | Rent, eviction, landlord-tenant disputes, habitability, deposits, mortgages, foreclosure, neighbors, housing subsidies. |
WORK | Employment contracts, wages, dismissal, discrimination at work, leave, workplace safety, severance, freelancers when the main issue is labor rights. |
PUBLIC_BENEFITS | Unemployment benefits, disability, pensions, welfare, public assistance, eligibility, reductions, sanctions, appeals on benefits. |
CRIMINAL_JUSTICE | Police, arrest, criminal charges, fines, prosecution, defense, victims' rights, probation, criminal procedure. |
CONSUMER_DEBT | Purchases, warranties, subscriptions, refunds, scams, debt collection, loans, bankruptcy, credit, repossession, consumer finance. |
CONTRACTS | Private civil agreements and breach/interpretation issues not better covered by work, housing, consumer, or business. |
IMMIGRATION | Visas, residence permits, asylum, citizenship, deportation, family migration, immigration status and related procedures. |
BUSINESS | Company formation, shareholder issues, commercial compliance, business operations, B2B disputes, self-employment when the main issue is business law. |
DATA_PRIVACY | Personal data, surveillance, GDPR/privacy rights, data deletion, consent, monitoring, platform data practices. |
INTELLECTUAL_PROPERTY | Copyright, trademark, patent, trade secrets, licensing, infringement, ownership of creative or technical works. |
CIVIL_RIGHTS | Discrimination outside employment/housing, free speech, due process, equal treatment, constitutional or human-rights style claims. |
INTERNATIONAL_CROSS_BORDER | Choice of law, jurisdiction, treaty-based questions, cross-border enforcement, multi-country disputes where cross-border law is central. |
OTHER | Genuinely legal but not covered above. |
Selection rules: use CONTRACTS only when the issue is mainly about a civil agreement and is not better captured by WORK, HOUSING, CONSUMER_DEBT, or BUSINESS. Use INTERNATIONAL_CROSS_BORDER only when the cross-border or jurisdictional aspect is central, not merely incidental.
| Split | N | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|---|
| Validation (best checkpoint) | 106 | 83.96% | — | — | — |
| Test (held-out, full input) | 107 | 83.18% | 84.89% | 83.18% | 83.18% |
| Test (user-only input) | 107 | 64.49% | 63.17% | 64.49% | 60.85% |
Joint pipeline (with modernbert-seeks_guidance, max length 4096, full input, test): legal 87.9% · topic 78.5% · joint 78.5%.
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
def serialize(messages, input_mode="full"):
lines = []
for msg in messages:
role = msg["role"]
if input_mode == "user" and role != "user":
continue
lines.append(f"{role.capitalize()}: {msg['content']}")
return "\n".join(lines)
def predict(model_id, text):
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
enc = tokenizer(text, truncation=True, max_length=4096, return_tensors="pt")
with torch.no_grad():
pred_id = model(**enc).logits.argmax(dim=-1).item()
label = model.config.id2label[str(pred_id)]
return label or "(non-guidance)"
conversation = [
{"role": "user", "content": "Can my landlord evict me without notice?"},
{"role": "assistant", "content": "Eviction rules depend on your jurisdiction..."},
{"role": "user", "content": "I'm in California on a month-to-month lease."},
]
topic = predict(
"AmirMohseni/modernbert-primary_topic",
serialize(conversation, input_mode="full"),
)
print(topic) # e.g. HOUSING
Pair with modernbert-seeks_guidance for the full routing pipeline (see that model card for a complete example).
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="AmirMohseni/modernbert-primary_topic",
)
text = "User: Can my landlord evict me?\nAssistant: Eviction rules depend on your jurisdiction...\nUser: I'm in California on a month-to-month lease."
print(classifier(text))
Use for: assigning a topic label to user queries already flagged as seeking legal guidance.
Do not use for: legal advice, or as a standalone filter for legal intent (use the seeks_guidance model first).
Caveats: silver labels from GPT-5.4; English only.
Dataset: AmirMohseni/WildChat-Legal-Classification-V2-Balanced
primary_topic (empty for non-guidance rows)| Setting | Value |
|---|---|
| Base model | answerdotai/ModernBERT-base |
| Input mode | Full conversation |
| Max length | 4096 |
| Learning rate | 1e-4 |
| Epochs | 8 |
| Effective batch size | 32 (8 × 4 grad accum) |
| Best checkpoint | Highest validation accuracy (83.96%) |