Downloads · 30 days
2.7K
52% of all-time downloads
fomofo/tap-ct-b-2d
tap-ct-b-2d is a machine learning model from fomofo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
Downloads · 30 days
2.7K
52% of all-time downloads
All-time downloads
5.3K
Public
Parameters
85.4M
342 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors342 MB · 100%
From the Hugging Face model README
TAP-CT is a suite of foundation models for computed tomography (CT) imaging, pretrained in a task-agnostic manner through an adaptation of DINOv2 for volumetric data. These models learn robust 3D representations from CT scans without requiring task-specific annotations.
This repository provides TAP-CT-B-2D, a Vision Transformer (ViT-Base) architecture pretrained on volumetric inputs with a spatial resolution of (224, 224) and a patch size of (16, 16). For inference on full-resolution CT volumes, a slice-wise inference approach can be employed to extract features across the entire scan.
Each TAP-CT model repository provides its own dedicated image processor and configuration file. To ensure proper preprocessing, it is recommended to instantiate the corresponding image processor using the AutoImageProcessor class from Hugging Face Transformers. This can be accomplished as follows:
from transformers import AutoImageProcessor
preprocessor = AutoImageProcessor.from_pretrained(
'fomofo/tap-ct-b-2d',
trust_remote_code=True
)
This approach automatically loads the appropriate processor and configuration for the selected TAP-CT model.
import torch
from transformers import AutoModel
# Load the model
model = AutoModel.from_pretrained('fomofo/tap-ct-b-2d', trust_remote_code=True)
# Prepare input (batch_size, channels, height, width)
x = torch.randn((16, 1, 224, 224))
# Forward pass
with torch.no_grad():
output = model.forward(x)
Recommended environment:
import numpy as np
import SimpleITK as sitk
import torch
from transformers import AutoModel, AutoImageProcessor
# Load the model
model = AutoModel.from_pretrained('fomofo/tap-ct-b-2d', trust_remote_code=True)
preprocessor = AutoImageProcessor.from_pretrained('fomofo/tap-ct-b-2d', trust_remote_code=True)
# Load image & set orientation to LPS
volume = sitk.ReadImage('/path/to/ct-scan.nii.gz')
volume = sitk.DICOMOrient(volume, 'LPS')
# Get array, expand to (B, C, D, H, W) and preprocess
array = sitk.GetArrayFromImage(volume)
array = np.expand_dims(array, axis=(0, 1))
x = preprocessor(array)['pixel_values']
# Forward pass
with torch.no_grad():
output = model.forward(x)
# OR
# Forward pass with slice-wise inference
from monai.inferers import SliceInferer
def predictor_fn(x):
# Reshape the patch tokens to resemble a 2D feature map
out = model(x, reshape=True)
return out.last_hidden_state
inferer = SliceInferer(
roi_size=[224, 224],
sw_batch_size=1,
spatial_dim=0
)
with torch.no_grad():
output = inferer(x, predictor_fn)
The model returns a BaseModelOutputWithPooling object from the transformers library. The output.pooler_output contains the pooled [CLS] token representation, while output.last_hidden_state contains the spatial patch token embeddings. To extract features from all intermediate transformer layers, pass output_hidden_states=True to the forward method.
(batch_size, 1, height, width)(16, 1, 224, 224) - batch of 16 CT slices at 224×224 resolutionIf you find this work useful, please cite:
@article{veenboer2025tapct,
title={TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models},
author={Veenboer, Tim and Yiasemis, George and Marcus, Eric and Van Veldhuizen, Vivien and Snoek, Cees G. M. and Teuwen, Jonas and Groot Lipman, Kevin B. W.},
journal={arXiv preprint arXiv:2512.00872},
year={2025}
}