Downloads ยท 30 days
0
krystv/PMA-VAE
PMA-VAE is a machine learning model from krystv. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A novel attention-free architecture for image generation, super-resolution, artifact removal, and artistic style transfer.
Downloads ยท 30 days
0
Access
Public
Updated Apr 30, 2026
Repo size
โ
Likes
0
Public
Click a slice to open those files.
.py64.3 KB ยท 48%
From the Hugging Face model README
A novel attention-free architecture for image generation, super-resolution, artifact removal, and artistic style transfer.
Image โ PixelUnshuffle stem โ MobileConv stages โ Parallel 2D Mamba blocks
โ Multi-scale latent (z_base H/16, z_detail H/8, z_style global)
โ Light parallel decoder with FiLM style modulation โ Reconstructed image
| Component | Choice | Why |
|---|---|---|
| Backbone | MobileConv + Parallel 2D Mamba | Fast, efficient, attention-free |
| Downsampling | PixelUnshuffle โ stride-2 conv | Lossless initial features |
| Upsampling | PixelShuffle (sub-pixel) | Mobile-friendly, no checkerboard |
| Latent | Multi-scale (base/detail/style) | Controllable, prevents collapse |
| Style control | FiLM conditioning | Lightweight, multiplicative |
| Global context | 4-dir cross-scan SSM | O(n) complexity, no attention |
| Local context | Depthwise separable conv + SE | Standard mobile building block |
| Training | Progressive resolution + KL warmup | Stable convergence |
| Loss | L1 + VGG + PatchGAN + edge + KL | Comprehensive quality |
z_base (structure), z_detail (texture), z_style (global style)| Config | Encoder | Decoder | Total | Target |
|---|---|---|---|---|
pmavae_tiny | 0.56M | 1.91M | 2.47M | Testing |
pmavae_small | 2.00M | 4.27M | 6.27M | Free Colab T4 |
pmavae_base | 5.18M | 9.83M | 15.01M | Colab Pro / better GPU |
z_base : H/16 ร W/16 ร 24-32 โ Structure, composition, objects
z_detail : H/8 ร W/8 ร 6-8 โ Texture, brush strokes, edges
z_style : 1 ร 1 ร 96-128 โ Global style vector
This separation enables:
z_style between imagesz_detail while keeping z_basez_detail while preserving structureLoss = L1 + 0.5 ร VGG_perceptual + 0.1 ร edge_sobel + ฮฒ ร KL_free_bits + ฮป ร PatchGAN
Phase 1: 256ร256 โ Learn structure
Phase 2: 384ร384 โ Refine texture
Phase 3: 512ร512 โ Full detail
Phase 4: FHD tiled โ High-resolution fine-tuning
The decoder is designed for mobile:
model.py โ Full PMA-VAE architecture (encoder, decoder, SSM blocks)losses.py โ Loss functions (VGG perceptual, PatchGAN, KL free bits, edge loss)train.py โ Training script with progressive resolution, checkpoint managementPMA_VAE_Colab_Training.ipynb โ Complete Colab notebookMIT