Downloads · 30 days
20
27% of all-time downloads
QuarkML/QaDiT-160
QaDiT-160 is a text-to-audio model from QuarkML. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center" <a href="https://github.com/sidharth72/QaDiT" <img src="https://img.shields.io/badge/GitHub-Repository-181717?logo=github&logoColor=white" alt="GitHub" </a <a href="https://www.quarkml.com" <img src=…
Downloads · 30 days
20
27% of all-time downloads
All-time downloads
74
Public
Parameters
159M
637 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors637 MB · 100%
From the Hugging Face model README
A ~159M-parameter latent Diffusion Transformer that turns a text caption into 10.24 s of 16 kHz mono audio: FLAN-T5 conditioning → DiT denoising of AudioLDM KL-VAE latents → VAE decode → HiFi-GAN vocoder.
Example: Prompt:
A small waterfall flows through a forest while insects buzz and birds sing.Output:
<audio controls src="https://cdn-uploads.huggingface.co/production/uploads/64054e5e0ab5e22719fc179f/_RGLUAPxdImTitbAAD0g6.wav"></audio>
| Piece | Choice |
|---|---|
| Backbone | DiT-B — depth 12, width 768, 12 heads, MLP ratio 4.0 (~159M) |
| Latent grid | [8, 256, 16] (channels × time × freq) |
| Patchify | 2×2 → 1024 tokens, fixed 2-D sincos positions |
| Text | FLAN-T5-large cross-attention every block + pooled text in adaLN-Zero |
| Train target | v-prediction |
| Noise schedule | cosine ᾱ, T = 1000 |
| Timestep sampling (train) | logit-normal |
| CFG | p_uncond = 0.1 train; default guidance 4.0 at sample |
| Sampler | DDIM, default 50 steps, η = 0 |
| Aux loss | REPA vs frozen AST features (train only) |
| Decode stack | cvssp/audioldm-s-full-v2 VAE + HiFi-GAN |

Heavy frozen models run once; the training loop never loads T5, the VAE, or the REPA encoder.

Only the DiT and its small glue layers receive gradients.




Forward process
$$ z_t = \sqrt{\bar{\alpha}_t},z_0
Network target
$$ v = \sqrt{\bar{\alpha}_t},\varepsilon
At sample time the DiT predicts v; we recover \hat{z}_0 and \hat{\varepsilon},
then step with DDIM (\eta = 0). CFG is applied in v-space with default
scale (s = 4.0). After DDIM, latents are divided by latent_scale ≈ 0.95035
before VAE decode — that whole chain is what model.generate() runs.
| Source | OpenSound/AudioCaps |
| Split | train · 45,178 clips after precompute |
| Clip length | 10.24 s @ 16 kHz |
| Cached fields | VAE latents, FLAN-T5 embeddings + mask, AST REPA targets |
latent_scale | 0.9503493000009796 (baked into config.json) |
AudioCaps is captioned environmental / everyday sound — not speech or music. Those domains are out of distribution for this checkpoint.
| Optimizer | AdamW, lr 1e-4, weight decay 0 |
| Steps | 23,999 (EMA exported) |
| Global batch | 256 (2 GPUs × microbatch 16 × grad accum 8) |
| EMA decay | 0.9999 |
| REPA | weight 0.5, decayed over 15k steps |
| AMP | yes |
Put W&B / TensorBoard screenshots (or exports) under assets/ using
the filenames below. Until then the images show as broken links on the Hub —
that is intentional so the slots are obvious.


pip install transformers diffusers soundfile sentencepiece
import soundfile as sf
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("QuarkML/QaDiT", trust_remote_code=True)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()
out = model.generate(
"A small waterfall flows through a forest while insects buzz and birds sing.",
num_inference_steps=200,
guidance_scale=16.0,
seed=0,
)
sf.write("sample.wav", out.audios[0], out.sampling_rate)
First generate downloads the frozen helpers this run was trained with:
google/flan-t5-large and the VAE + vocoder from cvssp/audioldm-s-full-v2.
output_type | Field | Content |
|---|---|---|
"np" (default) | audios | list of float32 numpy waveforms in [-1, 1] |
"pt" | audio_values | [B, num_samples] tensor |
"latent" | latents | [B, 8, 256, 16] scaled latents (skips VAE/vocoder) |
out = model.generate(
encoder_hidden_states=text_emb, # [B, 64, 1024]
encoder_attention_mask=text_mask, # [B, 64]
)
v = model(latents, timesteps, encoder_hidden_states, encoder_attention_mask).sample
Runs on CPU and CUDA. With dtype=torch.float16 or torch.bfloat16 the
DiT backbone runs in half precision; DDIM schedule math stays in float32.
Keep T5 / VAE / vocoder in float32 (FLAN-T5 overflows easily in fp16).
config.latent_scale (0.9503493) must match training precompute.
generate divides by it before VAE decode.repa_layer exists for REPA fine-tuning; inference ignores it.Intended use: education, reproduction of a small latent DiT audio stack, ablations, and a starting checkpoint for longer / wider training.
Not intended for: production SFX libraries, speech synthesis, music generation, or safety-critical audio.
Known limits of this checkpoint
This release is a research artifact, not a production host model. The architecture and sampling path are solid enough to build on; the ceiling is mostly data and compute:
Those levers will move quality more than inventing a new backbone for this size of model. Contributions and longer runs are welcome; treat this Hub page as a reproducible baseline, not a finished product.
@misc{qadit2026,
title = {QaDiT: A Text-to-Audio Latent Diffusion Transformer},
author = {Sidharth GN},
year = {2026},
note = {Research artifact. Weights and transformers remote-code loading.}
}
| Resource | |
|---|---|
| Dataset | OpenSound/AudioCaps |
| VAE / vocoder | cvssp/audioldm-s-full-v2 |
| Text encoder | google/flan-t5-large |