Downloads · 30 days
113
100% of all-time downloads
rsoohyun/SpatialBlock-3B-direct
SpatialBlock-3B-direct is a image-text-to-text model from rsoohyun. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains the SpatialBlock-3B-direct checkpoint from the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem.
Downloads · 30 days
113
100% of all-time downloads
All-time downloads
113
Public
Parameters
3.8B
7.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors7.5 GB · 100%
From the Hugging Face model README
This repository contains the SpatialBlock-3B-direct checkpoint from the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem.
It is a fine-tuned version of Qwen2.5-VL-3B-Instruct on the synthetic SpatialBlock-15k dataset. The model directly predicts answers to spatial reasoning tasks such as 3D-to-2D projection, viewpoint transformation, and structural combination.
For training details, evaluation results, and the companion “reason” model, please refer to the GitHub repository: https://github.com/rsoohyun/SpatialBlock.