Downloads · 30 days
0
RKB109/multimodal-document-retrieval-20260802-model
multimodal-document-retrieval-20260802-model is a visual document retrieval model from RKB109. Use it for the visual document retrieval task on the model card, and read the license before you ship it in a product. It is set up for custom. The card lists the license as mit.
This repository contains a small, transparent prototype model for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Downloads · 30 days
0
Access
Public
Updated Aug 2, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json6.3 KB · 51%
From the Hugging Face model README
This repository contains a small, transparent prototype model for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
The model combines per-label token weights with IDF-weighted evidence retrieval. It was generated for reproducible architecture demonstrations and does not call a hosted LLM.
visual-document-retrievaldocument-question-answeringimage-to-textfeature-extractionThe starter dataset contains synthetic textual modality descriptors, not sensitive scanned documents.
The dataset is synthetic and small. Do not use this model for consequential decisions without representative data, expert review, and production-grade evaluation.
The linked GitHub repository includes train.py, the exact dataset split,
evaluation code, and the model JSON format.