Downloads · 30 days
13
38% of all-time downloads
prolongvid/prolongvid_image_sft_7B
prolongvid_image_sft_7B is a image-text-to-text model from prolongvid. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
The ProLongVid-v1 models are 7B parameter models trained on ProLongViddata, based on our extended Qwen2.5 language model with a context window of 256K tokens.
Downloads · 30 days
13
38% of all-time downloads
All-time downloads
34
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
The ProLongVid-v1 models are 7B parameter models trained on ProLongVid_data, based on our extended Qwen2.5 language model with a context window of 256K tokens.
This prolongvid-image-sft-7B model is trained on the mid-training data and the single-image instruction tuning data of LLaVA-OneVision, based on our extended Qwen2.5 language model with a context window of 256K tokens.
We train the ProLongVid Video-LMMs based on this image-sft LMM.
@inproceedings{wang2025prolongvid,
title={ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning},
author={Wang, Rui and Li, Bohao and Dai, Xiyang and Yang, Jianwei and Chen, Yi-Ling and Xing, Zhen and Yang, Yifan and Chen, Dongdong and Qiu, Xipeng and Wu, Zuxuan and others},
booktitle={EMNLP},
year={2025}
}