Downloads · 30 days
0
MingliangLiang3/DynamiCS-ViT-B-16-DataComp-DFN
DynamiCS-ViT-B-16-DataComp-DFN is a zero-shot image classification model from MingliangLiang3. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for open_clip. The card lists the license as mit.
This repository hosts two OpenCLIP-compatible PyTorch checkpoints for DynamiCS, a dynamic cluster-based data sampling method for efficient and long-tail-aware vision-language pre-training.
Downloads · 30 days
0
Access
Public
Updated May 1, 2026
Repo size
14 GB
Likes
0
Public
Click a slice to open those files.
.json10.4 GB · 74%
From the Hugging Face model README
This repository hosts two OpenCLIP-compatible PyTorch checkpoints for DynamiCS, a dynamic cluster-based data sampling method for efficient and long-tail-aware vision-language pre-training.
The checkpoints correspond to the DataComp-DFN (130M) results reported in the DynamiCS project repository and paper draft, using a ViT-B/16 image encoder and the OpenCLIP text tower.
| File | Samples Seen @ Resolution | Tokens | ImageNet-1K | Let It Wag! | GPU-hours |
|---|---|---|---|---|---|
DynamiCS-ViT-B-16-DataComp-DFN-130M-1.28B.pt | 1.28B@112 + 128M@224 | 81 | 71.3 | 50.2 | 163 |
DynamiCS-ViT-B-16-DataComp-DFN-130M-2.56B.pt | 2.56B@112 + 128M@224 | 81 | 72.6 | 52.0 | 299 |
https://github.com/MingliangLiang3/DynamiCShttps://github.com/mlfoundations/open_clipDynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training.These checkpoints are intended for:
These files are stored as training checkpoints, not as Hub-native exported open_clip_pytorch_model.bin weights. They can be loaded with the DynamiCS/OpenCLIP codebase using open_clip.load_checkpoint, which extracts the state_dict automatically when needed.
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-16')
open_clip.load_checkpoint(model, '/path/to/DynamiCS-ViT-B-16-DataComp-DFN-130M-2.56B.pt')
tokenizer = open_clip.get_tokenizer('ViT-B-16')
model.eval()
The checkpoints were trained on a DataComp-DFN subset derived from DataComp-Large and filtered with DFN. In the project paper, the accessible subset is described as approximately 130M image-text pairs after accounting for unavailable or expired URLs.
DynamiCS computes per-sample sampling probabilities from semantic image clusters built with:
The exact web-scale training shards are not redistributed in this repository.
The training pipeline is based on OpenCLIP and the DynamiCS extensions in the GitHub repository.
50k0.70alpha = 0.250% of the accessible dataset per epochViT-B/1632112x112224x224amp_bf162 nodes x 4 H100 GPUs (8 GPUs total)1.28B@112 + 128M@224: lower-cost DynamiCS checkpoint2.56B@112 + 128M@224: longer-training DynamiCS checkpointThe primary reported metrics for these checkpoints are zero-shot top-1 classification on:
| Checkpoint | ImageNet-1K | Let It Wag! |
|---|---|---|
DynamiCS-ViT-B-16-DataComp-DFN-130M-1.28B.pt | 71.3 | 50.2 |
DynamiCS-ViT-B-16-DataComp-DFN-130M-2.56B.pt | 72.6 | 52.0 |
These results are taken from the project repository and accompanying paper draft.
The underlying code repository is released under the MIT License. Model users are responsible for ensuring that their use and any redistribution of checkpoints comply with the terms, restrictions, and policies associated with the underlying training data and their deployment context.
@article{liang2026dynamics,
title={Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training},
author={Mingliang Liang and Zhuoran Liu and Arjen P. de Vries and Martha Larson},
journal={arXiv preprint arXiv:2604.27932},
year={2026}
}