Downloads · 30 days
9
24% of all-time downloads
zeromodels/tipsv2-g14
tipsv2-g14 is a zero-shot image classification model from zeromodels. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for zeromodels. The card lists the license as apache-2.0.
[](https://github.com/IMvision12/ZeroModels) [](https://huggingface.co/collections/zeromodels/tipsv2-6a8eadf49c908c77e74e47eb)
Downloads · 30 days
9
24% of all-time downloads
All-time downloads
37
Public
Repo size
6.1 GB
Likes
0
Public
Click a slice to open those files.
.h56.1 GB · 100%
From the Hugging Face model README
Paper: TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment (arXiv:2604.12012)
TIPSv2 (Google DeepMind) is a CLIP/SigLIP-style dual encoder: a DINOv2-style ViT vision tower with register tokens plus a bidirectional text tower, aligned with a temperature-scaled contrastive objective.
For more details on the model, please go to the upstream model card.
Pure-Keras 3 conversion of google/tipsv2-g14 for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX. The full model and both towers load from this single repo.
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from PIL import Image
import numpy as np
import keras
from zeromodels.models.tipsv2 import Tipsv2Model, Tipsv2Processor
model = Tipsv2Model.from_weights("zeromodels/tipsv2-g14")
processor = Tipsv2Processor.from_weights("zeromodels/tipsv2-g14")
image = Image.open("your_image.jpg").convert("RGB")
texts = ["a photo of a cat", "a photo of a dog", "a photo of a car"]
inputs = processor(text=texts, images=np.array(image))
out = model(inputs)
probs = keras.ops.softmax(out["logits_per_image"], axis=-1)
print(keras.ops.convert_to_numpy(probs)[0])
Towers only:
from zeromodels.models.tipsv2 import Tipsv2VisionModel, Tipsv2TextModel
vision = Tipsv2VisionModel.from_weights("zeromodels/tipsv2-g14")
text = Tipsv2TextModel.from_weights("zeromodels/tipsv2-g14")
All TIPSv2 variants load the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub |
|---|---|
tipsv2-b14 | zeromodels/tipsv2-b14 |
tipsv2-l14 | zeromodels/tipsv2-l14 |
tipsv2-so400m14 | zeromodels/tipsv2-so400m14 |
tipsv2-g14 | zeromodels/tipsv2-g14 |
KERAS_BACKEND before importing Keras / zeromodels.[0, 1] (no mean/std normalization); input resolution is 448.Tipsv2Model.from_weights("hf:google/tipsv2-g14").A huge thank you to the TIPSv2 authors (Google DeepMind) and the HF community.
License: Apache-2.0 (matches the upstream google/tipsv2-g14 checkpoint).