Downloads · 30 days
63
100% of all-time downloads
Ambarella/OWLViT
OWLViT is a machine learning model from Ambarella. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
Downloads · 30 days
63
100% of all-time downloads
All-time downloads
63
Public
Repo size
1 GB
Likes
0
Public
Click a slice to open those files.
.bin1.3 GB · 100%
From the Hugging Face model README

OWL-ViT extends CLIP-based vision–language models to perform open-vocabulary object detection by aligning image regions with textual descriptions, enabling zero-shot detection without task-specific training.
Original paper: Simple Open-Vocabulary Object Detection with Vision Transformers
This model uses the OWL-ViT v1 with a CLIP ViT-B/32 Transformer architecture as an image encoder and a masked self-attention Transformer as a text encoder, leveraging the vision–language alignment of CLIP to detect objects specified by arbitrary text queries. It is well suited for applications such as open-vocabulary detection, image search, and real-time visual understanding across diverse domains.
Model Configuration:
| Model | Device | compression | Model Link |
|---|---|---|---|
| OWLv1 CLIP ViT-B/32 Image Encoder | N1-655 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | N1-655 | Amba_optimized | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | N1-655 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Predictor | N1-655 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Image Encoder | X7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | X7 | Amba_optimized | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | X7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Predictor | X7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Image Encoder | CV7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV7 | Amba_optimized | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Predictor | CV7 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Image Encoder | CV72 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV72 | Amba_optimized | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV72 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Predictor | CV72 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Image Encoder | CV75 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV75 | Amba_optimized | Model_Link |
| OWLv1 CLIP ViT-B/32 Text Encoder | CV75 | Activation_fp16 | Model_Link |
| OWLv1 CLIP ViT-B/32 Predictor | CV75 | Activation_fp16 | Model_Link |