Downloads · 30 days
0
vidore/colqwen2-base
colqwen2-base is a machine learning model from vidore. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for colpali. The card lists the license as apache-2.0.
ColQwen is a model based on a novel model architecture and training strategy based on Vision Language Models (VLMs) to efficiently index documents from their visual features. It is a Qwen2-VL-2B extension that generat…
Downloads · 30 days
0
Access
Public
Updated Jun 5, 2025
Parameters
2.2B
8.8 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
ColQwen is a model based on a novel model architecture and training strategy based on Vision Language Models (VLMs) to efficiently index documents from their visual features. It is a Qwen2-VL-2B extension that generates ColBERT- style multi-vector representations of text and images. It was introduced in the paper ColPali: Efficient Document Retrieval with Vision Language Models and first released in this repository
This version is the untrained base version to guarantee deterministic projection layer initialization.
[!WARNING] This version should not be used: it is solely the base version useful for deterministic LoRA initialization.
If you use any datasets or models from this organization in your research, please cite the original dataset as follows:
@misc{faysse2024colpaliefficientdocumentretrieval,
title={ColPali: Efficient Document Retrieval with Vision Language Models},
author={Manuel Faysse and Hugues Sibille and Tony Wu and Bilel Omrani and Gautier Viaud and Céline Hudelot and Pierre Colombo},
year={2024},
eprint={2407.01449},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2407.01449},
}