Downloads · 30 days
15
11% of all-time downloads
yhx12/VideoSSR
VideoSSR is a machine learning model from yhx12. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
VideoSSR-8B is a multimodal large language model (MLLM) fine-tuned from Qwen-VL-8B-Instruct for enhanced video understanding. It is trained using a novel Video Self-Supervised Reinforcement Learning (VideoSSR) framewo…
Downloads · 30 days
15
11% of all-time downloads
All-time downloads
135
Public
Parameters
8.8B
17.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors17.5 GB · 100%
From the Hugging Face model README
VideoSSR-8B is a multimodal large language model (MLLM) fine-tuned from Qwen-VL-8B-Instruct for enhanced video understanding. It is trained using a novel Video Self-Supervised Reinforcement Learning (VideoSSR) framework, which generates its own high-quality training data directly from videos, eliminating the need for manual annotation.
Qwen-VL-8B-Instruct