Downloads · 30 days
19
56% of all-time downloads
huggingbrain/Dinov3d-Neuro
Dinov3d-Neuro is a image feature extraction model from huggingbrain. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as cc-by-nc-sa-4.0.
This model uses the DINOv3 architecture and self-supervised training objective, adapted to 3D volumetric input via the nrdg/dinov3d codebase, itself a fork of facebookresearch/dinov3. Dinov3D was trained from scratch…
Downloads · 30 days
19
56% of all-time downloads
All-time downloads
34
Public
Repo size
81.9 GB
Likes
1
Public
Click a slice to open those files.
.pth70.6 GB · 86%
From the Hugging Face model README
This model uses the DINOv3 architecture and self-supervised training objective, adapted to 3D volumetric input via the
nrdg/dinov3dcodebase, itself a fork offacebookresearch/dinov3. Dinov3D was trained from scratch on brain MRI data and contains no weights from Meta's DINOv3 checkpoints. It is not affiliated with, endorsed by, or sponsored by Meta.
A 3D Vision Transformer encoder pretrained with a DINOv3 self-supervised objective (DINO + iBOT + KoLeo) with Gram anchoring added in later training phases, and a final high resolution adaptation phase (In Progress!). The model was trained on single-channel brain MRI volumes from the FOMO300K dataset of various contrasts (T1w, T2w, FLAIR, ...). Given a preprocessed volume, the model returns a CLS token and a grid of patch tokens (792-dim each).
vit3d_base in dinov3d repo(D/16)·(H/16)·(W/16) patch tokens (no
register tokens). model.forward_features(x) returns a dict including
x_norm_clstoken (B×792) and x_norm_patchtokens (B×N×792).asagilmore/dinov3d (training/modeling code)Adapted code: This model uses code adapted from
facebookresearch/dinov3, which is licensed
under the DINOv3 License.
weights are original work trained from scratch on brain MRI and are licensed separately under CC BY-NC-SA 4.0.
Modifications made for 3D:
| Component | 2D DINOv3 | This model |
|---|---|---|
| Patch embedding | Conv2d, 16×16 | Conv3d, 16×16×16 (in_chans=1, single-channel MRI) |
| Positional encoding | 2-axis RoPE | 3-axis (D/H/W) axial RoPE |
| Register tokens | 4 | 0 |
| Global/local crops | 2D multi-crop | 3D multi-crop: 128^3 global / 48^3 local (base + Gram phases), later 192^3 / 64^3 (high-res phase); 8 local crops per global crop |
| Color augmentation | color jitter, RGB mean/std normalization | none - single-channel intensities are z-score normalized once during preprocessing instead |
| Objective terms | DINO + iBOT + Koleo + Gram anchoring | Same four terms, applied in stages - see Training Details |
Data augmentation has also been reworked with domain specific augmentations for brain MRI, more details on data augmentation can be found in the code repo.
Evaluation and downstream adaptation is still in progress, check back later to see more details.
This model is distributed as a raw PyTorch checkpoint plus the code in this repo, as well as sharded checkpoints at the end of each training phase.
The teacher checkpoints can be loaded following the inference instructions in the repo, and the sharded checkpoints can be used to restart training for fine-tuning experiments.
The teacher checkpoints can be found in the eval folder. We include all checkpoints captured during training, but recommend using the latest one, unless doing experiments to evaluate performance over training iterations.
The ckpt directory contains the sharded checkpoints, which we include at the end of each training phase.
The layout of the checkpoints follows from the original dinov3 outputs, so their repo can be used as a rough reference for layout.
Dataset: FOMO-MRI/FOMO300K (Cerri et al., 2026)
We filtered the fomo300k dataset to include only single channel anatomical scans. The repo contains a fomo300k.json file listing all subjects used.
FOMO300K is distributed as NIfTI without co-registration or skull-stripping.
Spacingd, bilinear/trilinear interpolation, border padding),
applied at load time via InferenceAugmentation3d / the training data pipelineDivisiblePadd(..., mode="minimum"))NormalizeIntensityd(nonzero=True, channel_wise=True)), plus empty-signal filling (SignalFillEmptyd) - done once ahead of
time when building the preprocessed dataset (scripts/preproccess_fomo300k.py), not
per-forward-passAll training hyperparameters can be found in the config.yaml file in this HF repo.