Downloads · 30 days
11
15% of all-time downloads
robotflowlabs/clip-vit-large-patch14-int8
clip-vit-large-patch14-int8 is a zero-shot image classification model from robotflowlabs. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
OpenAI's CLIP vision encoder quantized to INT8 for real-time robotic perception. 4.5x smaller than the original — from 6.5 GB to 1.5 GB — while preserving zero-shot classification and visual grounding capabilities.
Downloads · 30 days
11
15% of all-time downloads
All-time downloads
73
Public
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.data1.2 GB · 80%
From the Hugging Face model README
OpenAI's CLIP vision encoder quantized to INT8 for real-time robotic perception. 4.5x smaller than the original — from 6.5 GB to 1.5 GB — while preserving zero-shot classification and visual grounding capabilities.
This model is part of the RobotFlowLabs model library, built for the ANIMA agentic robotics platform — a modular ROS2-native AI system designed to bring foundation model intelligence to real robots operating in the real world.
Large vision-language models like CLIP are essential for robotic scene understanding — identifying objects, understanding spatial relationships, and grounding natural language instructions to visual observations. But at 6.5 GB, the original CLIP ViT-L/14 is too heavy for edge deployment on devices like NVIDIA Jetson, Raspberry Pi, or embedded industrial controllers.
We quantized CLIP to INT8 and exported to ONNX so robots can run it in real-time, on-device, without cloud dependencies.
| Property | Value |
|---|---|
| Architecture | Vision Transformer (ViT-L/14) |
| Parameters | 304M (vision encoder) |
| Hidden Dimension | 1024 |
| Layers | 24 transformer blocks |
| Attention Heads | 16 |
| MLP Dimension | 4096 |
| Input Resolution | 224 × 224 |
| Patch Size | 14 × 14 |
| Tokens | 257 (256 patches + 1 CLS) |
| Original Model | openai/clip-vit-large-patch14 |
| License | MIT |
Quantized on an NVIDIA L4 24GB GPU using INT8 dynamic quantization with ONNX Runtime export.
| Metric | Original | INT8 Quantized | Change |
|---|---|---|---|
| Total Size | 6,529 MB | 1,451 MB | 4.5x smaller |
| INT8 Weights | — | 293 MB | Vision encoder only |
| ONNX Graph | — | 1,158 MB | Full model with optimizations |
| Quantization | FP32 | INT8 Dynamic | Per-tensor symmetric |
| Format | PyTorch | PyTorch INT8 + ONNX | Dual format |
clip-vit-large-patch14-int8/
├── model_int8.pt # 293 MB — INT8 quantized state dict
├── model.onnx # 2.0 MB — ONNX graph structure
├── model.onnx.data # 1.2 GB — ONNX external weights
├── config.json # Model configuration
├── preprocessor_config.json # Image preprocessing config
├── tokenizer_config.json # Text tokenizer config
└── README.md # This file
import torch
from transformers import CLIPModel, CLIPProcessor
# Load original architecture
model = CLIPModel.from_pretrained("openai/clip-vit-large-patch14")
# Load INT8 quantized vision encoder weights
int8_state = torch.load("model_int8.pt", map_location="cuda", weights_only=True)
model.vision_model.load_state_dict(int8_state, strict=False)
# Run inference
processor = CLIPProcessor.from_pretrained("openai/clip-vit-large-patch14")
inputs = processor(images=image, text=["a robot arm", "a table", "a cup"], return_tensors="pt")
outputs = model(**inputs)
import onnxruntime as ort
import numpy as np
# GPU inference
session = ort.InferenceSession(
"model.onnx",
providers=["CUDAExecutionProvider", "CPUExecutionProvider"]
)
# Preprocess image to (1, 3, 224, 224) float32
pixel_values = preprocess(image) # Your preprocessing pipeline
outputs = session.run(None, {"pixel_values": pixel_values})
from forge.vision import VisionEncoderRegistry
# FORGE auto-detects INT8 weights and loads optimally
encoder = VisionEncoderRegistry.load("clip-vit-large-patch14-int8")
features = encoder(image_tensor) # (B, 257, 1024)
CLIP serves as the visual grounding backbone across multiple ANIMA modules:
ANIMA is a modular, ROS2-native agentic robotics platform developed by RobotFlowLabs. It combines 58 specialized AI modules — from perception and planning to manipulation and safety — into a unified system that enables robots to understand, reason, and act in unstructured real-world environments.
ANIMA modules run on edge hardware (Jetson Orin, industrial PCs) with real-time constraints. Every foundation model we deploy must be compressed without sacrificing the capabilities that make it useful. That's why we built FORGE — our distillation and compression pipeline — and why we're releasing optimized model variants publicly.
We believe the robotics community deserves production-ready models, not just research checkpoints.
Browse all optimized models at huggingface.co/robotflowlabs:
Original CLIP ViT-L/14 (FP32, 6.5 GB)
│
├─→ torch.quantization.quantize_dynamic (INT8)
│ └─→ model_int8.pt (293 MB)
│
└─→ torch.onnx.export (opset 18, GPU-traced)
└─→ model.onnx + model.onnx.data (1.2 GB)
nn.Linear layersopenai/clip-vit-large-patch14 by OpenAI@article{radford2021learning,
title={Learning Transferable Visual Models From Natural Language Supervision},
author={Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and others},
journal={International Conference on Machine Learning},
year={2021}
}
@misc{robotflowlabs2026anima,
title={ANIMA: Agentic Networked Intelligence for Modular Autonomy},
author={RobotFlowLabs},
year={2026},
url={https://huggingface.co/robotflowlabs}
}