Downloads · 30 days
1.3K
15% of all-time downloads
fomofo/tap-ct-b-3d
tap-ct-b-3d is a machine learning model from fomofo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
Downloads · 30 days
1.3K
15% of all-time downloads
All-time downloads
8.6K
Public
Parameters
87.1M
348 MB on disk
Likes
10
Trending 1
Click a slice to open those files.
.safetensors348 MB · 100%
From the Hugging Face model README
TAP-CT is a suite of foundation models for computed tomography (CT) imaging, pretrained in a task-agnostic manner through an adaptation of DINOv2 for volumetric data. These models learn robust 3D representations from CT scans without requiring task-specific annotations.
This repository provides TAP-CT-B-3D, a Vision Transformer (ViT-Base) architecture pretrained on volumetric inputs with a spatial resolution of (12, 224, 224) and a patch size of (4, 8, 8). For inference on full-resolution CT volumes, a sliding window approach can be employed to extract features across the entire scan.
Each TAP-CT model repository provides its own dedicated image processor and configuration file. To ensure proper preprocessing, it is recommended to instantiate the corresponding image processor using the AutoImageProcessor class from Hugging Face Transformers. This can be accomplished as follows:
from transformers import AutoImageProcessor
preprocessor = AutoImageProcessor.from_pretrained(
'fomofo/tap-ct-b-3d',
trust_remote_code=True
)
This approach automatically loads the appropriate processor and configuration for the selected TAP-CT model.
import torch
from transformers import AutoModel
# Load the model
model = AutoModel.from_pretrained('fomofo/tap-ct-b-3d', trust_remote_code=True)
# Prepare input (batch_size, channels, depth, height, width)
x = torch.randn((16, 1, 12, 224, 224))
# Forward pass
with torch.no_grad():
output = model.forward(x)
Recommended environment:
import numpy as np
import SimpleITK as sitk
import torch
from transformers import AutoModel, AutoImageProcessor
# Load the model
model = AutoModel.from_pretrained('fomofo/tap-ct-b-3d', trust_remote_code=True)
preprocessor = AutoImageProcessor.from_pretrained('fomofo/tap-ct-b-3d', trust_remote_code=True)
# Load image & set orientation to LPS
volume = sitk.ReadImage('/path/to/ct-scan.nii.gz')
volume = sitk.DICOMOrient(volume, 'LPS')
# Get array, expand to (B, C, D, H, W) and preprocess
array = sitk.GetArrayFromImage(volume)
array = np.expand_dims(array, axis=(0, 1))
x = preprocessor(array)['pixel_values']
# Forward pass
with torch.no_grad():
output = model.forward(x)
# OR
# Forward pass with sliding window
from monai.inferers import SlidingWindowInferer
def predictor_fn(x):
# Reshape the patch tokens to resemble a 3D feature map
out = model(x, reshape=True)
return out.last_hidden_state
inferer = SlidingWindowInferer(
roi_size=[12, 224, 224],
sw_batch_size=1,
overlap=0.75,
mode='gaussian'
)
with torch.no_grad():
output = inferer(x, predictor_fn)
The model returns a BaseModelOutputWithPooling object from the transformers library. The output.pooler_output contains the pooled [CLS] token representation, while output.last_hidden_state contains the spatial patch token embeddings. To extract features from all intermediate transformer layers, pass output_hidden_states=True to the forward method.
(batch_size, 1, depth, height, width)(16, 1, 12, 224, 224) - batch of 16 CT crops with 12 slices at 224×224 resolutionIf you find this work useful, please cite:
@article{veenboer2025tapct,
title={TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models},
author={Veenboer, Tim and Yiasemis, George and Marcus, Eric and Van Veldhuizen, Vivien and Snoek, Cees G. M. and Teuwen, Jonas and Groot Lipman, Kevin B. W.},
journal={arXiv preprint arXiv:2512.00872},
year={2025}
}