Downloads · 30 days
0
zeqianli/HowToStep-NSVA
HowToStep-NSVA is a machine learning model from zeqianli. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Jan 31, 2024
Repo size
85.7 MB
Likes
1
Public
Click a slice to open those files.
.pth85.7 MB · 100%
From the Hugging Face model README

NSVA is a lightweight Transformer-based architecture, where the narration or procedural steps are used as queries, to iteratively attend the video features, and output the alignability or optimal temporal windows.
[project page] [Arxiv] [GitHub]
We provide pre-trained models for HTM-Align and HT-Step. You can use these two models for reproducing our results, following our [code].