Downloads · 30 days
0
KRMayD/COD10K_CAM_5Way_Segmentation_Checkpoints
COD10K_CAM_5Way_Segmentation_Checkpoints is a machine learning model from KRMayD. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for open_clip.
This repository contains the five OpenCLIP ViT-B-32-quickgelu checkpoints used in one controlled COD10K CAM segmentation comparison.
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Repo size
5.4 GB
Likes
0
Public
Click a slice to open those files.
.pt5.4 GB · 100%
From the Hugging Face model README
This repository contains the five OpenCLIP ViT-B-32-quickgelu checkpoints
used in one controlled COD10K CAM segmentation comparison.
| File | Training method |
|---|---|
baseline_openai_clip_vit_b32_quickgelu.pt | Unmodified OpenAI CLIP baseline |
clip_posttrained_explicitneg_2886_seed0_final_model.pt | Native CLIP post-training on positive image-caption pairs |
cliprefine_posttrained_explicitneg_2886_seed0_final_model.pt | Native CLIP-Refine post-training on positive image-caption pairs |
gmpo_global_ind_explicitneg_2886_seed0_epoch_3.pt | GMPO with one frozen reference-derived IND four-way vector for all samples |
gmpo_sample_ind_explicitneg_2886_seed0_epoch_3.pt | GMPO with a frozen reference-derived IND four-way vector per sample |
All checkpoints are plain PyTorch/OpenCLIP state dictionaries compatible with
the ViT-B-32-quickgelu architecture.
filename, filename_neg, Caption, Caption_neg.T+) or absence (T-) sentence.There is a/an {animal} visually blended into its surroundings.vbeta=1.0, vvar=0.3, vlayer=8, seed 42.The exact paths, settings, data checksum, and metrics are included in
comparison_protocol.json and metrics_summary.json.
| Model | DSC | NSD |
|---|---|---|
| Baseline CLIP | 0.352689 | 0.375161 |
| CLIP post-training | 0.359849 | 0.383442 |
| CLIP-Refine post-training | 0.359409 | 0.383454 |
| GMPO Global-IND | 0.378582 | 0.402299 |
| GMPO Sample-IND | 0.356773 | 0.379253 |
The train split, starting checkpoint, number of epochs, seed, test images,
prompt, saliency settings, and metric code are fixed across the comparison.
GMPO necessarily uses the full four-tuple. CLIP and CLIP-Refine use only the
positive image-caption pair by their native definitions and retain their
native optimizer settings; see comparison_protocol.json for the exact
values.