Downloads · 30 days
172
71% of all-time downloads
s-nlp/enoki-openie-encoder
enoki-openie-encoder is a feature extraction model from s-nlp. Use it when you need embeddings to search or compare text. It is set up for transformers.
Downloads · 30 days
172
71% of all-time downloads
All-time downloads
243
Public
Parameters
395M
1.6 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README

Enoki OpenIE Encoder is an LLM-free fact extractor for English. It distills
Enoki's LLM-based OpenIE decomposition into a ModernBERT-large Iterative Grid
Labeling encoder that converts sentences into text-anchored
(subject, relation, object) triples. Like the atomic-fact decomposition stage
in FActScore-style factuality pipelines, it turns free-form text into facts
that can be verified independently, but performs the extraction with a single
encoder instead of a generative LLM.
The model is designed for Enoki's multi-level hallucination detection pipeline. Extracted facts can be checked against a reference context with a separate NLI or factuality verifier. Because the extracted facts remain anchored to tokens in the original sentence, an unsupported fact can be projected directly back to its source span. The same representation supports both claim-level verification and span-level localization without a separate LLM-based claim-to-text alignment step.
The encoder was trained on the EnokiQA dev split using incremental triples
produced by Enoki-LLM, making the training setup a form of fact-extractor
distillation.
Source code: s-nlp/Enoki
Paper: Enoki: Efficient Multi-Level Hallucination Detection
Training data: s-nlp/EnokiQA
pip install torch "transformers>=4.48,<5" nltk
from transformers import AutoModel
model = AutoModel.from_pretrained(
"s-nlp/enoki-openie-encoder",
trust_remote_code=True,
)
results = model.extract_triples(
"Apple acquired Beats Electronics for $3 billion in 2014.",
min_confidence=0.7,
)
print(results)
Example output:
[
{
"sentence": "Apple acquired Beats Electronics for $3 billion in 2014.",
"triples": [
{
"subject": "Apple",
"relation": "acquired",
"object": "Beats Electronics",
"confidence": 0.969,
},
{
"subject": "Apple",
"relation": "acquired for",
"object": "$3 billion",
"confidence": 0.934,
},
],
}
]
For several sentences, pass a list to model.extract_triples([...]).
python inference.py \
--model s-nlp/enoki-openie-encoder \
--text "Barack Obama was born in Honolulu." \
--min-confidence 0.7
Use the Enoki package to combine this encoder with NLI-based fact verification and span-level hallucination localization:
git clone https://github.com/s-nlp/Enoki.git
cd Enoki
poetry install --all-extras
from enoki import EnokiPipeline
context = "Apple acquired Beats Electronics in 2014 for $3 billion."
answer = "Apple acquired Beats Electronics in 2015 for $3 billion."
results = EnokiPipeline(
model="s-nlp/enoki-openie-encoder"
).detect(context=context, answer=answer)
print(results)
[
{
"span": "2015",
"start": 36,
"end": 40,
"fact": {
"subject": "Apple",
"predicate": "acquired in",
"object": "Beats Electronics 2015",
},
"probability": 0.9625,
}
]
min_confidence=0.7–0.8
when a smaller, higher-precision result set is preferred.trust_remote_code=True because the IGL architecture and
OpenIE decoder are custom Transformers code included in this repository.If you use Enoki in your research, please cite:
@misc{rykov2026enokiefficientmultilevelhallucination,
title = {Enoki: Efficient Multi-Level Hallucination Detection},
author = {Elisei Rykov and Timur Ionov and Nikolay Ivanov and Maksim Savkin and Maksim Makarenko and Alexander Panchenko and Vasily Konovalov and Julia Belikova},
year = {2026},
eprint = {2609.00581},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.00581},
}