Downloads · 30 days
247
1% of all-time downloads
VITA-MLLM/VITA-1.5
VITA-1.5 is a video-text-to-text model from VITA-MLLM. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product.
This repository contains the model of the paper VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.
Downloads · 30 days
247
1% of all-time downloads
All-time downloads
17.4K
Public
Parameters
8.3B
19.6 GB on disk
Likes
50
Public
Click a slice to open those files.
.safetensors16.6 GB · 85%
From the Hugging Face model README
This repository contains the model of the paper VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.