Downloads · 30 days
79
13% of all-time downloads
Diankun/Spatial-MLLM-v1.1-Instruct-820K
Spatial-MLLM-v1.1-Instruct-820K is a video-text-to-text model from Diankun. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This repository contains the model described in Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.
Downloads · 30 days
79
13% of all-time downloads
All-time downloads
624
Public
Parameters
5.6B
11.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors11.2 GB · 100%
From the Hugging Face model README
This repository contains the model described in Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.
Project page: https://diankun-wu.github.io/Spatial-MLLM/