Downloads · 30 days
0
THUdyh/Oryx-ViT
Oryx-ViT is a image feature extraction model from THUdyh. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The Oryx-ViT model is trained on 200M data and can seamlessly and efficiently process visual inputs with arbitrary spatial sizes and temporal lengths. It is described in the paper Oryx MLLM: On-Demand Spatial-Temporal…
Downloads · 30 days
0
Access
Public
Updated Mar 1, 2025
Repo size
893 MB
Likes
8
Public
Click a slice to open those files.
.pth893 MB · 100%
From the Hugging Face model README
The Oryx-ViT model is trained on 200M data and can seamlessly and efficiently process visual inputs with arbitrary spatial sizes and temporal lengths. It is described in the paper Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.
@article{liu2024oryx,
title={Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution},
author={Liu, Zuyan and Dong, Yuhao and Liu, Ziwei and Hu, Winston and Lu, Jiwen and Rao, Yongming},
journal={arXiv preprint arXiv:2409.12961},
year={2024}
}