Downloads · 30 days
56
71% of all-time downloads
RamonK/DistillPath-IS16-Virchow2
DistillPath-IS16-Virchow2 is a image feature extraction model from RamonK. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for timm. The card lists the license as cc-by-nc-4.0.
A 22M ViT-S/16 pathology tile encoder distilled from Virchow2 (632M ViT-H/14) into an ImageNet-21k pretrained ViT-S/16 student using backbone-token distillation on 6,000 public TCGA slides.
Downloads · 30 days
56
71% of all-time downloads
All-time downloads
79
Public
Parameters
21.7M
86.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors86.7 MB · 100%
From the Hugging Face model README
A 22M ViT-S/16 pathology tile encoder distilled from Virchow2 (632M ViT-H/14) into an ImageNet-21k pretrained ViT-S/16 student using backbone-token distillation on 6,000 public TCGA slides.
This is the ImageNet-initialized variant. For the kaiko-initialized variant, see DistillPath-KS16-Virchow2.
Paper: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance Ramon Kaspar, Andrey Ignatov, Valentina Boeva. ETH Zurich. Published at the ECCV 2026 Workshop on Medical Foundation Models and Benchmarks (MedFM-Bench).
| Property | Value |
|---|---|
| Architecture | ViT-S/16 (vit_small_patch16_224 in timm) |
| Parameters | 21.7M |
| Feature dimension | 384 |
| Input size | 224 x 224 |
| Normalization | mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225) |
| Student initialization | ImageNet-21k ViT-S/16 |
| Teacher | Virchow2 (632M, ViT-H/14, CC BY-NC-ND 4.0) |
| Training data | 6,000 TCGA H&E whole-slide images, 32 cohorts |
| Training steps | 50,000 (batch size 256) |
| Benchmark | DistillPath-IS16-Virchow2 | IN21K baseline | Virchow2 teacher |
|---|---|---|---|
| EVA mean (7 tasks) | 0.763 | 0.729 | 0.810 |
| HEST mean (9 tasks) | 0.358 | 0.311 | 0.398 |
| PLISM score | 0.490 | 0.383 | 0.447 |
See the paper for per-task results.
Load directly from the Hub with timm:
import timm
model = timm.create_model(
"hf_hub:RamonK/DistillPath-IS16-Virchow2",
pretrained=True,
num_classes=0,
)
model.eval()
Or load manually:
import timm
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
model = timm.create_model("vit_small_patch16_224", pretrained=False, num_classes=0)
path = hf_hub_download("RamonK/DistillPath-IS16-Virchow2", "model.safetensors")
state_dict = load_file(path)
model.load_state_dict(state_dict, strict=True)
model.eval()
This model uses ImageNet normalization: mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225).
from torchvision import transforms
transform = transforms.Compose([
transforms.Resize(224),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])
The recipe reads only the teacher's final class and patch tokens (no teacher pretraining heads required):
Full details in the paper and the DistillPath repository.
This model is released under the Creative Commons Attribution-NonCommercial 4.0 International License. The Virchow2 teacher is released under CC BY-NC-ND 4.0. The distillation process used Virchow2 only to generate supervisory outputs; the released student contains no Virchow2 weights. This model is intended solely for non-commercial academic research.
@inproceedings{kaspar2026distillpath,
title = {DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance},
author = {Kaspar, Ramon and Ignatov, Andrey and Boeva, Valentina},
booktitle = {Medical Foundation Models and Benchmarks (MedFM-Bench), ECCV 2026},
year = {2026}
}