Downloads · 30 days
48
22% of all-time downloads
Pranshurs/groundcheck-modernbert
groundcheck-modernbert is a text classification model from Pranshurs. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
A ~150M-parameter encoder that checks whether a RAG answer is supported by the source it was given. Given an answer and a source (and optionally the question), it returns grounded or hallucinated with P(grounded). It…
Downloads · 30 days
48
22% of all-time downloads
All-time downloads
215
Public
Parameters
150M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors598 MB · 99%
From the Hugging Face model README
A ~150M-parameter encoder that checks whether a RAG answer is supported by the source it
was given. Given an answer and a source (and optionally the question), it returns
grounded or hallucinated with P(grounded). It runs on CPU.
0 = grounded, 1 = hallucinated, and P(grounded) = softmax(logits)[0].| Recommended revision | 998cec35563d6b90947409d1c7510adac7f7c80c (the groundcheck-rag package pins this) |
| Weights committed in | 0c7dd0636c0d58b5e3865db8f1deee5a5acfeb6c (later commits change only this card) |
model.safetensors SHA-256 | 9ec331ba6d8a9d93236dd72b239df518b07e61241323876f8aafc223d959461a |
| Base model | answerdotai/ModernBERT-base (Apache-2.0) |
Hashes for every file are in the repository's MODEL_PROVENANCE.json.
python -m eval.provenance --verify re-checks them.
All measured on CPU against the pinned weights, using test sets rebuilt from pinned public
datasets. F1 is for the hallucinated class, and the 95% CIs come from a 2,000-resample
bootstrap.
| Suite | n | max_length 512 (training protocol) | max_length 2048 (groundcheck-rag default) |
|---|---|---|---|
| RAGTruth test (first 2,500 of 2,700 responses) | 2,500 | F1 0.682 [0.658, 0.705], acc 0.746 | F1 0.696 [0.672, 0.718], acc 0.758 |
| VitaminC test | 2,000 | acc 0.850, F1 0.845 | identical (all inputs are short) |
| One-fact flips caught (regenerated holdout) | 500 | 78.0% | 87.2% |
| Same answers unflipped, kept grounded | 500 | 76.0% | 74.0% |
The RAGTruth paper (Niu et al., 2024; tabulated in LettuceDetect, 2025) reports F1 0.634 for a zero-shot GPT-4-turbo prompt judge. That figure comes from a different protocol: a prompted judge scored on all 2,700 test responses. It was not run head-to-head with GroundCheck, so it's there for orientation and doesn't support a "beats GPT-4" claim.
Apple M1, CPU, 4 torch threads, single requests after warm-up, at max_length 2048:
| Pair length | p50 |
|---|---|
| ≤ 128 tokens | 39 ms |
| 129–512 tokens | 172 ms |
| 513–2,048 tokens | 349 ms |
| > 2,048 tokens | ~1.45 s |
Other hardware will differ. Measure with python -m bench.latency.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
name, rev = "Pranshurs/groundcheck-modernbert", "998cec35563d6b90947409d1c7510adac7f7c80c"
tok = AutoTokenizer.from_pretrained(name, revision=rev)
model = AutoModelForSequenceClassification.from_pretrained(name, revision=rev).eval()
source = "France's capital and largest city is Paris."
answer = "Paris is the capital of France."
enc = tok(source, answer, truncation="only_first", max_length=512, return_tensors="pt")
with torch.no_grad():
grounded = torch.softmax(model(**enc).logits, dim=-1)[0, 0].item()
print("grounded" if grounded >= 0.5 else "hallucinated", round(grounded, 3))
Or use the library (pip install "groundcheck-rag[model]"), which pins the revision,
truncates only the source, and raises an error rather than substituting a heuristic when the
model can't load:
from groundcheck import GroundCheck
print(GroundCheck().check(source="...", answer="..."))
Fine-tuned from answerdotai/ModernBERT-base (pinned at 8949b909) as a sequence-pair
classifier: the premise is the optional question plus the source, and the hypothesis is the
answer. The training run was on a single Kaggle P100 with torch 2.4.1 and transformers 4.49.0:
3 epochs, batch 16, learning rate 2e-5, linear schedule, warmup 0.06, weight decay 0.01,
seed 42, fp16, sequence length 512.
The 28,500 training rows break down as:
wandb/RAGTruth-processed @ eb4f4b9d), with the question dropped
on ~50% of rows.tals/vitaminc @ be6febb7): SUPPORTS → grounded; REFUTES and NOT
ENOUGH INFO → hallucinated.There is no LLM-generated augmentation and no private data. The recipe and data builders
are in the repository's training/ directory.
Use it as a post-generation check in RAG pipelines: flag answers the retrieved source doesn't support, so they can be routed to review, regeneration or a stronger checker. It checks support against the provided source only and isn't a world-knowledge fact-checker.
hallucinated verdicts on long RAG
answers are false alarms. Scores aren't calibrated probabilities; tune the threshold on
your own data.| What | Terms |
|---|---|
| These model weights | MIT |
Base model, answerdotai/ModernBERT-base | Apache-2.0 |
| The GroundCheck code | Apache-2.0 |
| Training data | Keeps its upstream terms, which the MIT license on the weights does not change |
The upstream data terms (detailed in the repository's DATA_LICENSES.md):
Whether a dataset's non-commercial terms extend to a model trained on it is legally unsettled. Review the data terms before any commercial use. This card is not legal advice.