Downloads · 30 days
65
48% of all-time downloads
ddz16/CamDistill-8B
CamDistill-8B is a video-text-to-text model from ddz16. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Camera-movement understanding model trained with Camera Token Distillation on top of Qwen/Qwen3-VL-8B-Instruct. A lightweight Camera Token Module learns geometry-aware camera tokens (distilled from VGGT) and injects t…
Downloads · 30 days
65
48% of all-time downloads
All-time downloads
136
Public
Parameters
9B
36.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.4 GB · 100%
From the Hugging Face model README
Camera-movement understanding model trained with Camera Token Distillation on top of
Qwen/Qwen3-VL-8B-Instruct. A lightweight Camera Token Module learns geometry-aware camera
tokens (distilled from VGGT) and injects them into the language model. Given a video, it outputs
structured JSON describing every camera-movement segment.
⚠️ This model cannot be loaded with plain 🤗 Transformers. It contains an extra Camera Token Module and a patched forward pass. Loading it as a standard
Qwen3VLForConditionalGenerationwould silently drop those weights and produce incorrect results. Use the CamDistill repo, which registers the required custom model type through a plugin.
Clone the CamDistill repo, then run (camera tokens are generated internally — no online VGGT required):
python camera_movement_sft/infer_single.py \
--model ddz16/CamDistill-8B \
--video /path/to/video.mp4 \
--variant camdistill
See the repo's README for environment setup and batch evaluation.