Downloads · 30 days
9
3% of all-time downloads
RavenK/TAC-ViT-base-rgb
TAC-ViT-base-rgb is a feature extraction model from RavenK. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
This model is used for encoding RGB image into a dense feature.
Downloads · 30 days
9
3% of all-time downloads
All-time downloads
353
Public
Parameters
87.5M
700 MB on disk
Likes
0
Public
Click a slice to open those files.
.bin350 MB · 50%
From the Hugging Face model README
This model is used for encoding RGB image into a dense feature.
Caution, the model does not contain the last FC layer. So, the output features are not aligned with depth.
The model is pre-trained with RGB-D contrastive objectives, named TAC. Different from InfoNCE-based loss fuctions, TAC leverages the similarity between videos frames and estimate a similarity matrix as soft labels. The backbone of this version is ViT-B/32. The pre-training is conducted on a new unified RGB-D database, UniRGBD. The main purpose of this work is depth representation. So, the RGB encoder is just a side model.
@ARTICLE{10288539,
author={He, Zongtao and Wang, Liuyi and Dang, Ronghao and Li, Shu and Yan, Qingqing and Liu, Chengju and Chen, Qijun},
journal={IEEE Transactions on Circuits and Systems for Video Technology},
title={Learning Depth Representation From RGB-D Videos by Time-Aware Contrastive Pre-Training},
year={2024},
volume={34},
number={6},
pages={4143-4158},
doi={10.1109/TCSVT.2023.3326373}}