Downloads · 30 days
0
ra7an/mmvtg
mmvtg is a other model from ra7an. Use it for the other task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
This repository hosts the official model checkpoints accompanying our paper:
Downloads · 30 days
0
Access
Public
Updated Jul 5, 2026
Repo size
174 MB
Likes
0
Public
Click a slice to open those files.
.ckpt174 MB · 100%
From the Hugging Face model README
This repository hosts the official model checkpoints accompanying our paper:
Cross-Attention-Based Intelligent Video Temporal Grounding for CCTV-Based Crime Identification in Smart Cities
The released checkpoints are intended to support research reproducibility, benchmarking, and future work in multi-modal video understanding and intelligent surveillance.
This work extends the excellent Moment-DETR framework with a reproducible workflow tailored for Multi-Modal Video Temporal Grounding (MMVTG) in CCTV crime investigation scenarios.
Our contributions include:
This repository contains only the released model checkpoints.
The accompanying GitHub repository contains:
| Checkpoint | Filename | Purpose |
|---|---|---|
| ASR Pretrained | mmvtg-pretrained-asr.ckpt | Weakly-supervised pretraining using ASR captions |
| QVHighlights Fine-tuned | mmvtg-finetuned-qvhighlights.ckpt | Fine-tuned model for downstream MMVTG tasks |
| Best Validation Model | mmvtg-best-val.ckpt | Recommended checkpoint for evaluation, inference, and deployment |
The released checkpoints correspond to the following pipeline:
ASR Caption Pretraining
│
▼
Fine-tuning on QVHighlights
│
▼
Validation
│
▼
Best Checkpoint Selection
│
▼
Inference
│
▼
Grounded Inference using Retrieved FIR Reports
These checkpoints are released for:
This work extends the Moment-DETR framework.
The retrieval component is performed only during inference by augmenting the input query with retrieved FIR reports to provide additional contextual information.
This repository does not implement a complete Retrieval-Augmented Generation (RAG) architecture or an end-to-end retrieval-training pipeline.
This Hugging Face repository contains only the released model checkpoints.
The complete project is organized as:
If you use these checkpoints in your research, please cite our paper.
BibTeX will be added once the final publication metadata becomes available.
% Citation coming soon
This work builds upon the excellent open-source implementation of Moment-DETR developed by Lei et al.
We gratefully acknowledge the original authors for making their implementation publicly available and for providing the foundation upon which this work was developed.
| Resource | Link |
|---|---|
| Research Companion Repository | https://github.com/RA7AN/MMVTG |
| Research Paper | https://ieeexplore.ieee.org/document/11548957 |
| Original Moment-DETR Repository | https://github.com/jayleicn/moment_detr |
These checkpoints are released under the Apache License 2.0. See the accompanying GitHub repository for full licensing details and usage terms.