Downloads · 30 days
53
13% of all-time downloads
Voxel51/fc-clip
fc-clip is a image segmentation model from Voxel51. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
FC-CLIP is an open-vocabulary panoptic segmentation model that pairs a frozen ConvNeXt-Large CLIP backbone with a lightweight Mask2Former decoder. It achieves strong zero-shot performance without requiring separate sp…
Downloads · 30 days
53
13% of all-time downloads
All-time downloads
402
Public
Parameters
20.7M
82.9 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors82.9 MB · 100%
From the Hugging Face model README
FC-CLIP is an open-vocabulary panoptic segmentation model that pairs a frozen ConvNeXt-Large CLIP backbone with a lightweight Mask2Former decoder. It achieves strong zero-shot performance without requiring separate specialist models for things vs. stuff.
This repository hosts the COCO Panoptic checkpoint uploaded to HuggingFace Hub by Claude, for testing use with the FiftyOne Model Zoo.
Paper: "A Simple Framework for Open-Vocabulary Segmentation and Detection"
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, Xiaolong Wang.
CVPR 2023 · arxiv 2311.15539
Original code: bytedance/fc-clip — Apache 2.0 license
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("neerajaabhyankar/fc-clip", trust_remote_code=True)
model.eval()
# Preprocess: RGB uint8 numpy/PIL → normalised tensor
pixel_values = model.preprocess_image(your_pil_image) # [1, 3, H, W]
with torch.no_grad():
results = model(pixel_values)
panoptic_seg, segments_info = results[0]
# panoptic_seg: int32 tensor [H, W] — pixel → segment id
# segments_info: list[{"id", "category_id", "isthing"}]
results = model(pixel_values, class_names=["cat", "dog", "sky", "grass"])
| Component | Detail |
|---|---|
| Backbone | OpenCLIP ConvNeXt-Large (convnext_large_d_320, laion2b_s29b_b131k_ft_soup), frozen |
| Pixel decoder | 6-layer Multi-Scale Deformable Attention encoder + 1-level FPN |
| Transformer decoder | 5-layer Mask2Former cross-attention decoder, 250 queries |
| Text classification | VILD 14-template ensemble + geometric in-vocab/out-vocab blending |
| Classes | 133 COCO panoptic (80 things + 53 stuff) |
torch torchvision transformers open_clip_torch safetensors
No detectron2 required — the model is self-contained.
Apache 2.0 (same as original bytedance/fc-clip)