Downloads · 30 days
16
1% of all-time downloads
LinWeizheDragon/FLMR
FLMR is a feature extraction model from LinWeizheDragon. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
FLMR is an open-source model for multimodal knowledge retrieval. It is a transformer-based model that uses a combination of text and image inputs to retrieve relevant documents from a large corpus.
Downloads · 30 days
16
1% of all-time downloads
All-time downloads
1.5K
Public
Parameters
207M
828 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors828 MB · 100%
From the Hugging Face model README
FLMR is an open-source model for multimodal knowledge retrieval. It is a transformer-based model that uses a combination of text and image inputs to retrieve relevant documents from a large corpus.
This model can be used directly to retrieve documents from a large corpus using a combination of text and image input queries. The retrieval usage can be found in the official implementation.
This model can be used combined with language models to create a retrieval-augmented language model. The use for Knowledge-based VQA can be found in RAVQA
For details of training, indexing, and performing retrieval, please refer to here.
The model is pre-trained on
For details on the dataset split and conversion process, please refer to the paper Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering.
The processed datasets are:
The model is evaluated on OKVQA, Infoseek, and FVQA.
Please find the evaluation results in the paper.
BibTeX:
@inproceedings{
lin2023finegrained,
title={Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering},
author={Weizhe Lin and Jinghong Chen and Jingbiao Mei and Alexandru Coca and Bill Byrne},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
url={https://openreview.net/forum?id=IWWWulAX7g}
}