Downloads · 30 days
12
14% of all-time downloads
Hevven/UFVideo-7B
UFVideo-7B is a machine learning model from Hevven. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository provides the complete code and datasets for UFVideo, a Video LLM that flexibly unifies general question answering, video object referring, video segmentation, and temporal video grounding to achieve mu…
Downloads · 30 days
12
14% of all-time downloads
All-time downloads
85
Public
Parameters
8.7B
17.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors17.3 GB · 100%
From the Hugging Face model README
This repository provides the complete code and datasets for UFVideo, a Video LLM that flexibly unifies general question answering, video object referring, video segmentation, and temporal video grounding to achieve multi-grained video understanding.
<!-- <p align="center"><img width="750" src="https://raw.githubusercontent.com/Heven-Pan/UFVideo/refs/heads/main/figs/overall_tasks.png"></p> -->First, clone the repository and navigate to the project folder.
git clone https://github.com/Heven-Pan/UFVideo
cd UFVideo
Then, install the requirement packages.
conda create -n UFVideo python=3.10.14
conda activate UFVideo
# our cuda version is 'cu124'
pip install -r requirements.txt
# other versions have no been verified
pip install flash-attn --no-build-isolation
Please kindly cite our paper if you find this project helpful.
@article{pan2025ufvideo,
title={UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models},
author={Pan, Hewen and Wei, Cong and Liang, Dashuang and Huang, Zepeng and Gao, Pengfei and Zhou, Ziqi and Xue, Lulu and Yan, Pengfei and Wei, Xiaoming and Li, Minghui and others},
journal={arXiv preprint arXiv:2512.11336},
year={2025}
}