Downloads · 30 days
591
60% of all-time downloads
WangYipu2002/CroPond-7B
CroPond-7B is a image-to-text model from WangYipu2002. Use it when you need a caption or text from an image. The card lists the license as mit.
[](https://arxiv.org/abs/2512.04686) [](https://github.com/WangYipu2002/CrossPoint) [](https://huggingface.co/datasets/WangYipu2002/CrossPoint-Bench)
Downloads · 30 days
591
60% of all-time downloads
All-time downloads
993
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
CroPond-7B is a vision-language model specialized in cross-view point correspondence. Built upon Qwen2.5-VL-7B-Instruct and mainly trained on the CrossPoint-378K dataset, CroPond achieves state-of-the-art performance on cross-view correspondence tasks.
For detailed evaluation instructions, please visit the GitHub repository.
@article{wang2025crosspoint,
title={Towards Cross-View Point Correspondence in Vision-Language Models},
author={Wang, Yipu and Ji, Yuheng and Liu, Yuyang and Zhou, Enshen and Yang, Ziqiang and Tian, Yuxuan and Qin, Ziheng and Liu, Yue and Tan, Huajie and Chi, Cheng and Ma, Zhiyuan and Zeng, Daniel Dajun and Zheng, Xiaolong},
journal={arXiv preprint arXiv:2512.04686},
year={2025}
}