Downloads · 30 days
26
53% of all-time downloads
ddz16/CamInject-4B
CamInject-4B is a video-text-to-text model from ddz16. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Camera-movement understanding model that injects frozen VGGT camera tokens into Qwen/Qwen3-VL-4B-Instruct. Given a video, it outputs structured JSON describing every camera-movement segment.
Downloads · 30 days
26
53% of all-time downloads
All-time downloads
49
Public
Parameters
4.4B
9.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.7 GB · 100%
From the Hugging Face model README
Camera-movement understanding model that injects frozen VGGT camera tokens into
Qwen/Qwen3-VL-4B-Instruct. Given a video, it outputs structured JSON describing every
camera-movement segment.
⚠️ This model cannot be loaded with plain 🤗 Transformers. It requires a custom model type (registered via a plugin) and runs VGGT online to produce camera tokens. Loading it as a standard
Qwen3VLForConditionalGenerationwould not work correctly. Use the CamDistill repo.
Clone the CamDistill repo and clone VGGT-Omega (set VGGT_OMEGA_REPO, see the repo's
setup). CamInject runs VGGT online during inference:
VGGT_TEACHER_TYPE=vggt_omega \
python camera_movement_sft/infer_single.py \
--model ddz16/CamInject-4B \
--video /path/to/video.mp4 \
--variant caminject
See the repo's README for environment setup and batch evaluation.