Downloads · 30 days
345
100% of all-time downloads
timm/vit_base_patch16_sapiens2.fb
vit_base_patch16_sapiens2.fb is a image feature extraction model from timm. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for timm. The card lists the license as other.
NOTE: This is a native timm (EVA) remap of facebook/sapiens2-pretrain-0.1b. Checkpoint keys have been converted to timm naming; the weights have not been fine-tuned. The original Sapiens2 License applies. The upstream…
Downloads · 30 days
345
100% of all-time downloads
All-time downloads
345
Public
Parameters
114M
456 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors456 MB · 100%
From the Hugging Face model README
NOTE: This is a native timm (EVA) remap of facebook/sapiens2-pretrain-0.1b. Checkpoint keys have been converted to timm naming; the weights have not been fine-tuned. The original Sapiens2 License applies. The upstream model card is reproduced below with timm usage instructions.
Sapiens2 is a family of high-resolution vision transformers pretrained on 1 billion human images — designed for human-centric tasks such as pose estimation, body-part segmentation, surface normals, and pointmaps.
This repository contains the 0.1B parameter pretrained backbone. It produces dense per-patch features suitable for fine-tuning downstream task heads.
model.safetensorsUse a timm version that includes Sapiens2 support.
import torch
import timm
from PIL import Image
device = "cuda" if torch.cuda.is_available() else "cpu"
model = timm.create_model(
"hf-hub:timm/vit_base_patch16_sapiens2.fb", pretrained=True, use_naflex=False,
).eval().to(device)
data_config = timm.data.resolve_model_data_config(model)
transform = timm.data.create_transform(**data_config, is_training=False)
image = Image.open("image.jpg").convert("RGB")
x = transform(image).unsqueeze(0).to(device)
with torch.inference_mode():
tokens = model.forward_features(x)
cls_features = tokens[:, 0]
patch_features = tokens[:, model.num_prefix_tokens:] # exclude CLS and register tokens
model(x) uses CLS-token pooling by default, matching the original Sapiens2 convention.
Pass global_pool="avg" to create_model for average pooling over patch tokens.
| Field | Value |
|---|---|
| Architecture | Sapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) |
| Parameters | 0.114 B |
| FLOPs | 0.342 T |
| Embedding dim | 768 |
| Layers | 12 |
| Attention heads | 12 |
| Pretraining resolution | 1024 × 768 (H × W) |
| Patch size | 16 |
| Pretraining data | 1B human images |
| Model | Params | FLOPs | Embed dim | Layers | Heads |
|---|---|---|---|---|---|
| Sapiens2-0.1B (this) | 0.114 B | 0.342 T | 768 | 12 | 12 |
| Sapiens2-0.4B | 0.398 B | 1.260 T | 1024 | 24 | 16 |
| Sapiens2-0.8B | 0.818 B | 2.592 T | 1280 | 32 | 16 |
| Sapiens2-1B | 1.462 B | 4.715 T | 1536 | 40 | 24 |
| Sapiens2-1B-4K | 1.607 B | — | 1536 | 40 | 24 |
| Sapiens2-5B | 5.071 B | 15.722 T | 2432 | 56 | 32 |
See the Sapiens2 Collection for all variants and downstream task checkpoints (pose, segmentation, normals, pointmaps).
Released under the Sapiens2 License.
@article{khirodkarsapiens2,
title={Sapiens2},
author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
journal={arXiv preprint arXiv:2604.21681},
year={2026}
}