Downloads · 30 days
15
1% of all-time downloads
AbstractPhil/vae-lyra-xl-adaptive-cantor
vae-lyra-xl-adaptive-cantor is a machine learning model from AbstractPhil. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Multi-modal Variational Autoencoder for SDXL text embedding transformation using adaptive Cantor fractal fusion with learned alpha (visibility) and beta (capacity) parameters.
Downloads · 30 days
15
1% of all-time downloads
All-time downloads
1.1K
Public
Repo size
166 GB
Likes
1
Public
Click a slice to open those files.
.pt1.6 GB · 100%
From the Hugging Face model README
Multi-modal Variational Autoencoder for SDXL text embedding transformation using adaptive Cantor fractal fusion with learned alpha (visibility) and beta (capacity) parameters.
Fuses CLIP-L, CLIP-G, and decoupled T5-XL scales into a unified latent space.
Alpha (Visibility):
Beta (Capacity):
This model outputs both CLIP embeddings needed for SDXL:
clip_l: [batch, 77, 768] → text_encoder outputclip_g: [batch, 77, 1280] → text_encoder_2 outputT5 information is encoded into the latent space and influences both CLIP outputs through learned binding weights.
from geovocab2.train.model.vae.vae_lyra import MultiModalVAE, MultiModalVAEConfig
from huggingface_hub import hf_hub_download
import torch
# Download model
model_path = hf_hub_download(
repo_id="AbstractPhil/vae-lyra-xl-adaptive-cantor",
filename="model.pt"
)
# Load checkpoint
checkpoint = torch.load(model_path)
# Create model
config = MultiModalVAEConfig(
modality_dims={
"clip_l": 768,
"clip_g": 1280,
"t5_xl_l": 2048,
"t5_xl_g": 2048
},
modality_seq_lens={
"clip_l": 77,
"clip_g": 77,
"t5_xl_l": 512,
"t5_xl_g": 512
},
binding_config={
"clip_l": {"t5_xl_l": 0.3},
"clip_g": {"t5_xl_g": 0.3},
"t5_xl_l": {},
"t5_xl_g": {}
},
latent_dim=2048,
fusion_strategy="adaptive_cantor",
cantor_depth=8,
cantor_local_window=3
)
model = MultiModalVAE(config)
model.load_state_dict(checkpoint['model_state_dict'])
model.eval()
# Use model - train on all four modalities
inputs = {
"clip_l": clip_l_embeddings, # [batch, 77, 768]
"clip_g": clip_g_embeddings, # [batch, 77, 1280]
"t5_xl_l": t5_xl_l_embeddings, # [batch, 512, 2048]
"t5_xl_g": t5_xl_g_embeddings # [batch, 512, 2048]
}
# For SDXL inference - only decode CLIP outputs
recons, mu, logvar, per_mod_mus = model(inputs, target_modalities=["clip_l", "clip_g"])
# Use recons["clip_l"] and recons["clip_g"] with SDXL
@software{vae_lyra_adaptive_cantor_2025,
author = {AbstractPhil},
title = {VAE Lyra: Adaptive Cantor Multi-Modal Variational Autoencoder},
year = {2025},
url = {https://huggingface.co/AbstractPhil/vae-lyra-xl-adaptive-cantor}
}