Downloads · 30 days
464
63% of all-time downloads
KRLabsOrg/lettucedect-v2-taxonomy-head
lettucedect-v2-taxonomy-head is a machine learning model from KRLabsOrg. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
464
63% of all-time downloads
All-time downloads
741
Public
Parameters
307M
649 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors614 MB · 95%
From the Hugging Face model README

LettuceDetect goes agentic — span-level hallucination detection across RAG, code, and tool output.
lettucedect-v2-taxonomy-head types a hallucinated span — it does not find spans. It is a
label-conditioned mmBERT-base bi-encoder that, given a span a binary detector already
located, assigns a hallucination category and subcategory by embedding the span and
taking the nearest taxonomy-label description (cosine). Paired with the binary encoder
lettucedect-v2-mmbert-base, it forms a fully-encoder typed detector — detection + typing
at encoder cost, no generative model.
lettucedect-v2-mmbert-base) first, then this head types each span.lettucedect-v2-qwen-2b.from lettucedetect.models.inference import HallucinationDetector
det = HallucinationDetector(
method="transformer",
model_path="KRLabsOrg/lettucedect-v2-mmbert-base", # binary detector (finds spans)
taxonomy_head="KRLabsOrg/lettucedect-v2-taxonomy-head", # this head (types them)
)
spans = det.predict(context=[context], question=question, answer=answer, output_format="spans")
# [{"start": ..., "end": ..., "text": "...", "category": "contradiction", "subcategory": "numerical"}]
A label-conditioned bi-encoder. A shared mmBERT-base encoder embeds (i) a span — the
mean-pooled answer tokens of the input answer [CTX] context — and (ii) each taxonomy label
as the mean-pooled embedding of its "name: description" text. Training minimizes cross-entropy
between the span embedding and the label embeddings (temperature-scaled cosine similarity),
jointly over the category and subcategory label sets, so a span is typed by its nearest label
description. Because labels enter only as text, the same head scores any label in the taxonomy.
Trained on the per-span category/subcategory annotations of the unified code + prose hallucination benchmark — SWE-bench coding-agent traces, developer tool output, ACL / README / Wikipedia, plus RAGTruth and 14-language PsiloQA.
lettucedect-v2-qwen-2b types in a
single pass and is higher (typed-F1 0.585 / 0.468); this cascade is the option when you want
typed spans from a small, fast encoder-only stack.@misc{kovács2026documentgroundingspanlevelhallucination,
title={Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents},
author={Ádám Kovács and Bowei He and Xue Liu and István Boros and Szilveszter Tóth and Gábor Recski},
year={2026},
eprint={2607.00895},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.00895},
}