Downloads · 30 days
19
23% of all-time downloads
VideoSearchR1/activitynet-stage2
activitynet-stage2 is a video-text-to-text model from VideoSearchR1. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This is the Stage 2 VideoSearch-R1 checkpoint trained for ActivityNet, presented in the paper VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement.
Downloads · 30 days
19
23% of all-time downloads
All-time downloads
81
Public
Parameters
2.5B
4.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.9 GB · 100%
From the Hugging Face model README
This is the Stage 2 VideoSearch-R1 checkpoint trained for ActivityNet, presented in the paper VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement.
Stage 2 starts from the ActivityNet Stage 1 checkpoint and optimizes iterative retrieval and temporal grounding behavior with the VideoSearch-R1 training pipeline.
Use with the VideoSearch-R1 codebase:
bash scripts/data_construct/download_preextracted_data.bash activitynet
EVAL_GPUS=0 bash scripts/inference/inference.bash activitynet --checkpoint VideoSearchR1/activitynet-stage2
@inproceedings{lee2026videosearchr1,
title = {VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement},
author = {Lee, Seohyun and Choi, Seoung and Ko, Dohwan and Kim, Jongha and Kim, Hyunwoo J.},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}