Downloads · 30 days
40
48% of all-time downloads
ddz16/CamSFT-4B
CamSFT-4B is a video-text-to-text model from ddz16. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Camera-movement understanding model, supervised fine-tuned from Qwen/Qwen3-VL-4B-Instruct. Given a video, it outputs structured JSON describing every camera-movement segment — time span, basic-movement type / directio…
Downloads · 30 days
40
48% of all-time downloads
All-time downloads
84
Public
Parameters
4.4B
9.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.7 GB · 100%
From the Hugging Face model README
Camera-movement understanding model, supervised fine-tuned from Qwen/Qwen3-VL-4B-Instruct.
Given a video, it outputs structured JSON describing every camera-movement segment — time span,
basic-movement type / direction / speed, and special techniques.
The CamDistill repo provides a one-line entry point that applies the official prompt and the exact video settings used for training and evaluation:
python camera_movement_sft/infer_single.py \
--model ddz16/CamSFT-4B \
--video /path/to/video.mp4
CamSFT is a standard Qwen3-VL model, so it can be loaded directly:
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
model = Qwen3VLForConditionalGeneration.from_pretrained(
"ddz16/CamSFT-4B", dtype="bfloat16", device_map="auto"
)
processor = AutoProcessor.from_pretrained("ddz16/CamSFT-4B")
The exact system/user prompt and the video preprocessing (fps, max frames, resolution) are provided in the CamDistill repo; using them is required to reproduce the paper's results.