Downloads · 30 days
65
1% of all-time downloads
LiheYoung/SigLIP-HD
SigLIP-HD is a image feature extraction model from LiheYoung. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
SigLIP-HD is a vision encoder fine-tuned from SigLIP 2-So400m/16-512px with fine-to-coarse supervision.
Downloads · 30 days
65
1% of all-time downloads
All-time downloads
10.8K
Public
Parameters
429M
1.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
From the Hugging Face model README
SigLIP-HD is a vision encoder fine-tuned from SigLIP 2-So400m/16-512px with fine-to-coarse supervision.
SigLIP-HD exhibits better performance than SigLIP 2 in MLLMs, especially for OCR scenarios.
This repository contains only the vision encoder (no text tower). It is a drop-in replacement for the SigLIP 2 vision tower: identical architecture and I/O. To use it in an MLLM, keep your existing SigLIP 2 pipeline and only change the vision-tower path to this checkpoint.
import torch
from PIL import Image
from transformers import SiglipVisionModel, AutoImageProcessor
model = SiglipVisionModel.from_pretrained("LiheYoung/SigLIP-HD").eval()
processor = AutoImageProcessor.from_pretrained("LiheYoung/SigLIP-HD")
image = Image.open("example.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
features = model(**inputs, output_hidden_states=True).hidden_states[-1] # (1, 1024, 1152)
@inproceedings{sigliphd,
title={SigLIP-HD by Fine-to-Coarse Supervision},
author={Yang, Lihe and Zhao, Zhen and Zhao, Hengshuang},
booktitle={ICLR},
year={2026}
}
This work is built upon SigLIP 2. We sincerely thank the authors for open-sourcing their excellent vision encoder.