Downloads · 30 days
2
3% of all-time downloads
hwh-datascience/divae-tokenizer-ahn-448
divae-tokenizer-ahn-448 is a machine learning model from hwh-datascience. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for terratorch. The card lists the license as apache-2.0.
A DiVAE (Diffusion VQ-VAE) tokenizer for Fused AHN6 DSM + DTM elevation (2-channel, float32) trained on Dutch national geospatial data. Part of the TerraVision-NL project.
Downloads · 30 days
2
3% of all-time downloads
All-time downloads
80
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.pt1.1 GB · 50%
From the Hugging Face model README
A DiVAE (Diffusion VQ-VAE) tokenizer for Fused AHN6 DSM + DTM elevation (2-channel, float32) trained on Dutch national geospatial data. Part of the TerraVision-NL project.
| Component | Value |
|---|---|
| Encoder | ViT-B (vit_b_enc) |
| Decoder | Patched UNet (unet_patched) |
| Quantizer | FSQ (codebook: 8-8-8-6-5, vocab: 15,360) |
| Image size | 448×448 px |
| Patch size | 16×16 px |
| Token grid | 28×28 = 784 tokens per image |
| Input channels | 2 (Digital Surface Model, Digital Terrain Model) |
| Latent dim | 5 |
All TerraVision tokenizers produce the same spatial window in 448 pixels, regardless of the underlying raster resolution. This ensures token grids are spatially aligned across modalities for cross-modal pretraining.
Input data should be normalized before encoding:
See config.json for exact normalization parameters.
import torch
from huggingface_hub import hf_hub_download
from terratorch.models.backbones.terramind.tokenizer.vqvae import DiVAE
# Download weights and config
weights_path = hf_hub_download(repo_id="YOUR_REPO_ID", filename="tokenizer.pt")
# Instantiate model
tokenizer = DiVAE(
image_size=448,
patch_size=16,
n_channels=2,
enc_type="vit_b_enc",
dec_type="unet_patched",
quant_type="fsq",
codebook_size="8-8-8-6-5",
latent_dim=5,
post_mlp=True,
norm_codes=True,
)
# Load weights
state_dict = torch.load(weights_path, map_location="cpu")
tokenizer.load_state_dict(state_dict)
tokenizer.eval()
# Encode: image → tokens
x = torch.randn(1, 2, 448, 448)
quant, code_loss, tokens = tokenizer.encode(x)
print(tokens.shape) # (1, 28, 28)
# Decode: tokens → reconstruction (diffusion sampling)
recon = tokenizer(x, timesteps=50)
Trained with the TerraVision-NL codebase using DiVAE (diffusion-based VQ-VAE) following the TerraMind paper methodology (Section 8.1).
ahn-best-epoch-0002.ckpt