Downloads · 30 days
17
15% of all-time downloads
newmindai/bge-m3-HD
bge-m3-HD is a token classification model from newmindai. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
bge-m3-HD is a multilingual embedding model fine-tuned for Turkish hallucination detection in Retrieval-Augmented Generation (RAG) systems. Based on the BAAI/bge-m3 model, this variant demonstrates exceptional perform…
Downloads · 30 days
17
15% of all-time downloads
All-time downloads
110
Public
Parameters
567M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
bge-m3-HD is a multilingual embedding model fine-tuned for Turkish hallucination detection in Retrieval-Augmented Generation (RAG) systems. Based on the BAAI/bge-m3 model, this variant demonstrates exceptional performance, particularly in Data2txt tasks, achieving the highest F1-score (84.62%) among all evaluated models in that domain.
This model leverages multilingual pretraining capabilities to provide strong cross-lingual understanding while being fine-tuned specifically for Turkish hallucination detection. It is part of the Turk-LettuceDetect project, which adapts the LettuceDetect framework for Turkish language applications.
Data2txt Task (Exceptional - Best in Class):
QA Task:
Summary Task:
Token-Level Task Performance:
This model is designed for:
The model was fine-tuned on RAGTruth-TR, a Turkish translation of the RAGTruth benchmark dataset:
The model was evaluated on the RAGTruth-TR test set across three task types:
Use this model when:
Consider alternatives when:
from lettucedetect import TransformerDetector
# Load the model
detector = TransformerDetector.from_pretrained(
"newmindai/bge-m3-HD"
)
# Detect hallucinations
context = "Your source document text..."
question = "Your question..."
answer = "Generated answer text..."
result = detector.detect(
context=context,
question=question,
answer=answer
)
# Access token-level predictions
hallucinated_tokens = result.get_hallucinated_spans()
If you use this model, please cite:
@misc{taş2025turklettucedetecthallucinationdetectionmodels,
title={Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications},
author={Selva Taş and Mahmut El Huseyni and Özay Ezerceli and Reyhan Bayraktar and Fatma Betül Terzioğlu},
year={2025},
eprint={2509.17671},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.17671},
}
@misc{chen2024bgem3embeddingmultilingualmultifunctionality,
title={BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
author={Jianlv Chen and Shitao Xiao and Peitian Zhang and Kun Luo and Defu Lian and Zheng Liu},
year={2024},
eprint={2402.03216},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2402.03216},
}
For questions or issues, please open an issue on the project repository.