Downloads · 30 days
26
15% of all-time downloads
garlandchou/V-Reflection
V-Reflection is a visual question answering model from garlandchou. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
V-Reflection is a Qwen2.5-VL-based multimodal model that turns the MLLM into an active interrogator via a think-then-look visual reflection mechanism: fixed-length latent visual reasoning before answering.
Downloads · 30 days
26
15% of all-time downloads
All-time downloads
176
Public
Parameters
921K
16.9 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors16.8 GB · 99%
From the Hugging Face model README
V-Reflection is a Qwen2.5-VL-based multimodal model that turns the MLLM into an active interrogator via a think-then-look visual reflection mechanism: fixed-length latent visual reasoning before answering.
| Resource | Link |
|---|---|
| Paper | arxiv.org/abs/2604.03307 |
| Project page | idea-research.github.io/V-Reflection |
| Code & scripts | github.com/IDEA-Research/V-Reflection |
This Hub repo includes model weights (sharded Safetensors) and pre-formatted LVR training annotations used in the paper project:
| File(s) | Role |
|---|---|
*.safetensors, config.json, tokenizer assets | Fine-tuned V-Reflection checkpoint |
meta_data_lvr_sft_stage1.json | Meta config for default SROIE + DUDE mix |
viscot_sroie_dude_lvr_formatted.json | SROIE + DUDE subset |
viscot_363k_lvr_formatted.json | Full Visual CoT–style 363K split |
Download images for Visual CoT / listed datasets from their official sources (see the code repo data section).
Loading and evaluation rely on the QwenWithLVR implementation and scripts in the GitHub repository (training, packing, and benchmarks are documented there).
README.md (environment, data layout).hf_hub_download from their evaluation scripts).EVAL_CHECKPOINT_PATH (or the training script checkpoint args) to the unpacked folder.Example evaluation entrypoint (from upstream docs):
bash scripts_release/evaluation/evaluation_7b_stage2.sh
| Benchmark | V-Reflection | Qwen2.5-VL-7B |
|---|---|---|
| MMVP | 72.3 | 66.7 |
| BLINK | 56.4 | 54.5 |
| V* | 81.7 | 78.5 |
| HRBench-4K | 72.6 | 68.0 |
| HRBench-8K | 66.3 | 63.8 |
| MME-RealWorld-Lite | 53.9 | 45.8 |
Apache-2.0 (same as the codebase). See the license file in the GitHub repository.