Downloads · 30 days
20
3% of all-time downloads
whybe-choi/kovre-stage1
kovre-stage1 is a sentence similarity model from whybe-choi. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
<p align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/655eeb5532537bcc8d7460ab/VR-ElTS3dohPp-deYSsgt.png" alt="KoVRE cover" width="600" / </p
Downloads · 30 days
20
3% of all-time downloads
All-time downloads
714
Public
Parameters
2.1B
4.3 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors4.3 GB · 100%
From the Hugging Face model README
KoVRE-Stage1 is a 2B-parameter single-vector embedding model for Korean visual document retrieval. It directly matches text queries against rendered document-page images and serves as the contrastive-learning checkpoint used before the knowledge-distillation stage of KoVRE.
The model is initialized from Qwen/Qwen3-VL-Embedding-2B and trained on 708,729 Korean and English query-page pairs using positive-aware hard-negative mining, self-guide filtering, hardness weighting, in-batch negatives, and Matryoshka Representation Learning.
For the strongest Korean VDR performance, use the final whybe-choi/kovre checkpoint, which further applies reranker-based knowledge distillation.
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3-VL-Embedding-2B |
| Parameters | 2B |
| Training stage | Stage 1: Contrastive learning |
| Representation | Single vector |
| Full embedding dimension | 2,048 |
| Matryoshka training dimensions | 128, 256, 512, 768, 1,024, 2,048 |
| Input modalities | Text query and document-page image |
| Similarity function | Cosine similarity |
| Pooling | Last-token pooling |
| Maximum image resolution used in training | 1,280 visual tokens (approximately 1.3M pixels) |
| Training languages | Korean and English |
[!NOTE] For production use or the best reported Korean benchmark performance, we recommend
whybe-choi/kovre.
KoVRE-Stage1 is intended for:
# pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("whybe-choi/kovre-stage1")
queries = [
"2024년 정보보호 예산은 얼마인가요?",
"재생에너지 발전량 추이를 보여주는 표",
]
document_pages = [
"pages/page_1.png",
"pages/page_2.png",
"pages/page_3.png",
]
query_embeddings = model.encode(
queries,
prompt="Find a document image that matches the given query.",
normalize_embeddings=True,
)
page_embeddings = model.encode(
document_pages,
normalize_embeddings=True,
)
scores = model.similarity(query_embeddings, page_embeddings)
print(scores)
For a smaller index, initialize the model with a Matryoshka dimension:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"whybe-choi/kovre-stage1",
truncate_dim=256,
)
The model was trained with the following instructions:
Find a document image that matches the given query.Represent the user's input. (the default instruction)KoVRE preserves the architecture of Qwen3-VL-Embedding. For lower-level inference with Transformers, use the Qwen3-VL-Embedding inference code and replace the model path with whybe-choi/kovre.
Recommended dependencies:
transformers>=4.57.0
qwen-vl-utils>=0.0.14
torch>=2.8.0
Find a document image that matches the given query.For Matryoshka inference, truncate the full embedding to the desired prefix dimension and L2-normalize it before computing similarity. The dimensions explicitly optimized during training were 128, 256, 512, 768, 1,024, and 2,048.
| Language | Query-page pairs |
|---|---|
| Korean | 406,945 |
| English | 301,784 |
| Total | 708,729 |
The Korean data combines the public training resource released with KoViDoRe and an additional private Korean collection. The English data combines five visual document retrieval resources covering reports, slides, tables, and other visually structured pages. English examples are included to help preserve the backbone's existing retrieval ability during Korean adaptation.
Seven hard negatives are mined per query using Qwen3-VL-Embedding-8B. Within each source dataset, document pages are ranked by cosine similarity after excluding known positives. Candidates scoring above 95% of the annotated positive score are removed to reduce false negatives, and query-positive pairs whose positive score does not exceed 0.3 are filtered out.
Training uses an InfoNCE objective over the paired positive, seven mined hard negatives, and in-batch negatives. The objective includes:
-0.1alpha = 2The full 2B model is trained for one epoch with bfloat16 mixed precision. GradCache is used to support a larger effective contrastive batch under the available GPU memory.
We evaluate KoVRE with nDCG@10 on two Korean visual document retrieval benchmarks:
AVG is the mean across the four KoViDoRe domains. OVR is the macro-average across those four domains and SDS KoPub VDR.
KoVRE is released under the Apache 2.0 License.
If you find KoVRE useful, please cite:
@article{choi2026kovre,
title={KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval},
author={Choi, Yongbin and Shim, Gyuho and Jang, Youngjoon},
journal={arXiv preprint arXiv:2608.01389},
year={2026}
}