Skip to content

VITA-MLLM

VITA-1.5

VITA-MLLM/VITA-1.5

VITA-1.5 is a video-text-to-text model from VITA-MLLM. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product.

This repository contains the model of the paper VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Downloads · 30 days

247

1% of all-time downloads

All-time downloads

17.4K

Public

Parameters

8.3B

19.6 GB on disk

Likes

50

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors16.6 GB · 85%

At a glance

Task
Video-Text-to-Text
Model type
vita-Qwen2
Access
Public
Created
Dec 18, 2024
Updated
Jan 16, 2025
SHA
e42b47fa
Task
Video-Text-to-Text
Type
vita-Qwen2
Created
Dec 18, 2024
Updated
Jan 16, 2025