Downloads · 30 days
21
48% of all-time downloads
cqqq/GSR-8B
GSR-8B is a image-text-to-text model from cqqq. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
GSR-8B is the full-parameter fine-tuned model released with Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection (CVPR 2026). It is trained on paired…
Downloads · 30 days
21
48% of all-time downloads
All-time downloads
44
Public
Parameters
770K
17.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.5 GB · 100%
From the Hugging Face model README
GSR-8B is the full-parameter fine-tuned model released with Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection (CVPR 2026). It is trained on paired top-view and side-view X-ray images from GSXray to perform geometric and semantic cross-modal reasoning.
Qwen/Qwen3-VL-8B-Instruct0c351dd01ed87e9c1b53cbc748cba10e6187ff3bcqqq/GSXray, 44,019 paired-image examples56f45e826f828e44fcdca6a1a5a854d4b71f6ec7The exact revision used by the original local base-model download was not recorded. The revision above is the immutable public reproduction pin selected from the upstream history; it is not presented as a recovered download record.
The model adds paired special tokens for top, side, think, conclusion,
answer, bbox, and label. For reproducible preprocessing and prompting,
load the processor and chat template distributed in this repository.
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
repo_id = "cqqq/GSR-8B"
model = Qwen3VLForConditionalGeneration.from_pretrained(
repo_id,
torch_dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(repo_id)
Pass the two X-ray views in the same top-view then side-view order used by the GSXray records. The public GSR code repository provides the complete prompt, training, serving, and evaluation entry points.
The preserved, path-free hyperparameters are in training_config.yaml; the
recorded final Trainer metrics are in training_summary.json. The release
contains inference weights only. Optimizer state, scheduler state, RNG state,
intermediate checkpoints, TensorBoard logs, and machine-local paths are
intentionally excluded.
This research model is intended for reproducible evaluation of dual-view X-ray reasoning. It can produce incorrect descriptions, locations, or conclusions and must not be treated as an autonomous security decision system. Results can depend on image acquisition conditions, prompts, decoding settings, and data distribution. Human review and application-specific validation are required.
checksums.sha256 covers every published file except the checksum file itself.
release_manifest.json records the closed release file set and provenance pins.
@inproceedings{peng2026gsr,
title = {Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection},
author = {Peng, Chuang and Tao, Renshuai and Ren, Zhongwei and Liu, Xianglong and Wei, Yunchao},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}
GSR-8B is a modified derivative of Qwen3-VL-8B-Instruct. See NOTICE for
copyright and upstream attribution.