Downloads · 30 days
0
AnuragBodkhe/contractguard-clause-classifier
contractguard-clause-classifier is a text classification model from AnuragBodkhe. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as other.
A fine-tuned Legal-BERT sequence classifier that labels a contract clause with one of 40 clause types. It is the first stage of ContractGuard AI, an open-source clause-level contract analysis platform — the clause typ…
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
438 MB
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
A fine-tuned Legal-BERT sequence classifier that labels a contract clause with one of 40 clause types. It is the first stage of ContractGuard AI, an open-source clause-level contract analysis platform — the clause type this model predicts conditions every downstream stage (risk scoring, retrieval, negotiation suggestions).
Not legal advice. This model produces a classification, not a legal conclusion. Its output is intended as decision support for a qualified reviewer, not a substitute for one.
| Base model | nlpaueb/legal-bert-base-uncased (110M parameters) |
| Task | Multi-class sequence classification, 40 clause types |
| Training data | CUAD + LEDGAR (union), 88,115 examples after validation |
| Version | 2.0 (active) — supersedes 1.0 (deprecated) |
| Max sequence length | 256 tokens |
| License | Other — see License |
This is version 2.0 of the classifier. It replaces v1.0, which was trained on LEDGAR alone and never activated in production — see Why v2.0, not v1.0 below.
Given the text of a single contract clause, predict which of 40 clause types
it is (e.g. TERMINATION, INDEMNIFICATION, NON_COMPETE,
LIMITATION_OF_LIABILITY). The predicted label and its confidence score feed
a downstream risk model, which conditions its own prediction on the clause
type, contract type, party role and jurisdiction.
In scope: English-language commercial contracts, segmented into individual clauses before inference.
Out of scope: whole-document classification, non-English contracts, and any of the six clause types listed in Coverage below, which the model was not trained to emit. It also does not detect a missing clause — absence is handled by a separate rule engine, since there is no text for a per-clause classifier to act on.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("AnuragBodkhe/contractguard-clause-classifier")
model = AutoModelForSequenceClassification.from_pretrained("AnuragBodkhe/contractguard-clause-classifier")
clause = "Either party may terminate this Agreement upon thirty (30) days' written notice."
inputs = tokenizer(clause, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
label_id = probs.argmax().item()
print(model.config.id2label[label_id], probs[0, label_id].item())
The label set and display names are also published in labels.py in the main
repository, which is the single source of truth the rest of the platform
mirrors — see Taxonomy below.
| Dataset | Role |
|---|---|
| CUAD | commercial contract clauses, 21 of 40 label types |
| LEDGAR | SEC-filed contract clauses, 28 of 40 label types |
91,144 raw examples were combined and reduced to 88,115 after validation, then split by document (not by clause) into train / validation / test to prevent leakage between splits:
| Split | Examples |
|---|---|
| Train | 61,483 |
| Validation | 13,151 |
| Test | 13,481 |
CUAD and LEDGAR each cover a different, overlapping subset of the 40-label taxonomy. Their union covers 34 of 40 (85%):
LEDGAR covers 28 / 40
CUAD covers 21 / 40
UNION 34 / 40
CUAD is what makes EXCLUSIVITY, LIABILITY, LICENSE_GRANT,
NON_COMPETE, NON_SOLICITATION and RENEWAL emittable at all — and
LIABILITY and NON_COMPETE are exactly the two types whose absence sank
v1.0 (see below).
Six labels have no training support in this release and the model cannot
produce them: DATA_PROTECTION, FORCE_MAJEURE, MORAL_RIGHTS,
PROBATION, SERVICE_COMMITMENT_BOND, SUBCONTRACTING. A label absent from
training data is a label the model cannot produce — closing this gap requires
adding labelled examples and retraining, not a configuration change.
| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Batch size | 16 × 2 gradient accumulation |
| Learning rate | 3e-5 |
| Max sequence length | 256 |
| Class weighting | sqrt-balanced |
| Precision | fp16 |
| Hardware | 1× RTX 3050 6 GB (laptop GPU), CUDA 12.6 |
| Wall clock | 59 minutes |
Measured on the held-out test split (13,481 examples, split by document):
| Metric | v1.0 (LEDGAR only) | v2.0 (CUAD + LEDGAR) |
|---|---|---|
| Accuracy | 0.8624 | 0.9168 |
| Macro F1 | 0.8592 | 0.8923 |
| Weighted F1 | 0.8618 | 0.9167 |
| Expected calibration error | 0.1136 | 0.0391 |
The calibration improvement is as consequential as the F1 gain: the downstream UI shows this model's confidence next to every prediction, and v1.0 was materially overconfident.
v1.0 scored a respectable 0.8592 macro F1 and was still never activated in
production, because LEDGAR's 28-label coverage disproportionately excluded
the clause types that carry risk. Adding CUAD lifted LIABILITY to F1 0.946
and NON_COMPETE to F1 0.660 — both previously unlearnable.
Aggregate test-set metrics can hide exactly the failure that matters, so activation was decided on a real document: a UK freelance agreement, analysed once with the deterministic baseline and once with v2.0.
| Baseline | v2.0 | |
|---|---|---|
| Overall risk score | 88 (CRITICAL) | 88 (CRITICAL) |
| Clause types | — | 10 of 11 identical to baseline |
| Risk findings | 11 | 11, none lost |
The one clause-type disagreement (RENEWAL vs. TERM_AND_DURATION) produced
the same downstream risk finding either way. Run against v1.0, the same test
lost two CRITICAL-severity findings — the failure mode aggregate metrics did
not surface.
v2.0 also declares the population it was trained on, and flags a real mismatch on this document:
fitted on United States (federal), United States – other state. This contract was submitted under United Kingdom.
fitted on COMMERCIAL, LICENSING, SERVICE, VENDOR contracts; this is a FREELANCE contract.
Both are true, and both are useful things for a reviewer to know before trusting the output.
The base model, nlpaueb/legal-bert-base-uncased, is licensed CC-BY-SA-4.0.
Training data licenses: CUAD (CC BY 4.0), LEDGAR (subject to the LexGLUE /
EDGAR terms). Review those upstream licenses before commercial use in
combination with this model's own license terms.
If you use this model, please cite the base model and the datasets it was fine-tuned on:
@article{chalkidis2020legalbert,
title={LEGAL-BERT: The Muppets straight out of Law School},
author={Chalkidis, Ilias and others},
year={2020}
}
@article{hendrycks2021cuad,
title={CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review},
author={Hendrycks, Dan and others},
year={2021}
}
This model is one of five in the ContractGuard AI pipeline (clause classification, risk classification, legal NER, retrieval embeddings, contract NLI). See the main repository for the full architecture, training results for all five models, and the taxonomy definitions this model uses.