Downloads · 30 days
0
salamnocap/ml-figs-ldm
ml-figs-ldm is a text-to-image model from salamnocap. Use it when you need an image from a text prompt. The card lists the license as cc-by-nc-4.0.
ML-FIGS-LDM is a Latent Diffusion Model (LDM) for generating educational figures. The AutoencoderKL is trained using a Text Perceptual Loss to reconstruct more readable text within the figures.
Downloads · 30 days
0
Access
Public
Updated May 20, 2025
Repo size
14.9 GB
Likes
0
Public
Click a slice to open those files.
.ckpt14.9 GB · 100%
From the Hugging Face model README
ML-FIGS-LDM is a Latent Diffusion Model (LDM) for generating educational figures. The AutoencoderKL is trained using a Text Perceptual Loss to reconstruct more readable text within the figures.
Note:
AutoencoderKL_SD: Stable Diffusion v1-4 autoencoder trained on LAION.AutoencoderKL_TPL: Autoencoders trained with Text Perceptual Loss (TPL).| Dataset | Method | PSNR ↑ | SSIM ↑ | FID ↓ | LPIPS ↓ | MSE ↓ | TPL ↓ |
|---|---|---|---|---|---|---|---|
| ML-Figs Test | AutoencoderKL_SD | 33.01 | 0.970 | 20.51 | 0.022 | 0.003 | 0.043 |
AutoencoderKL_TPL A | 30.71 | 0.954 | 16.13 | 0.056 | 0.002 | 0.017 | |
| ML-Figs + SciCap Test | AutoencoderKL_SD | 32.60 | 0.970 | 12.69 | 0.023 | 0.004 | 0.061 |
AutoencoderKL_TPL A | 29.94 | 0.954 | 9.235 | 0.057 | 0.003 | 0.028 | |
AutoencoderKL_TPL B | 31.47 | 0.979 | 6.256 | 0.016 | 0.001 | 0.010 |