Skip to content

Diankun

Spatial-MLLM-v1.1-Instruct-820K

Diankun/Spatial-MLLM-v1.1-Instruct-820K

Spatial-MLLM-v1.1-Instruct-820K is a video-text-to-text model from Diankun. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.

This repository contains the model described in Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

Downloads · 30 days

79

13% of all-time downloads

All-time downloads

624

Public

Parameters

5.6B

11.2 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors11.2 GB · 100%

At a glance

Task
Video-Text-to-Text
Library
transformers
License
mit
Model type
spatial-mllm
Access
Public
Created
Jan 5, 2026
Updated
Jan 17, 2026
SHA
89ac646b

Base models

Task
Video-Text-to-Text
Library
transformers
Type
spatial-mllm
License
mit
Created
Jan 5, 2026
Updated
Jan 17, 2026