Downloads · 30 days
4
9% of all-time downloads
dinhhung1508/SwiftAudio
SwiftAudio is a text-to-audio model from dinhhung1508. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for diffusers.
SwiftAudio: One-step Text-to-Audio Diffusion-Based Generation with an Audio-Free Distillation Technique
Downloads · 30 days
4
9% of all-time downloads
All-time downloads
47
Public
Parameters
860M
4.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.8 GB · 99%
From the Hugging Face model README
SwiftAudio: One-step Text-to-Audio Diffusion-Based Generation with an Audio-Free Distillation Technique
SwiftAudio is a one-step text-to-audio diffusion model. It distills a pretrained multi-step text-to-audio teacher into a one-step student using text captions only, without requiring paired audio data during distillation. The method adapts Variational Score Distillation (VSD) to the audio domain.
This repository contains the checkpoint used by the public SwiftAudio Gradio demo.
| Property | Value |
|---|---|
| Task | Text-to-audio generation |
| Output | Mono waveform, 16 kHz, approximately 10 seconds |
| Inference | One-step latent prediction |
| Text encoder | CLIP text encoder |
| Backbone | Conditional UNet |
| Audio decoder | Spectrogram VAE and neural vocoder |
| Distillation method | Variational Score Distillation |
Although the checkpoint uses Diffusers' StableDiffusionPipeline container,
SwiftAudio generates mel spectrograms rather than natural images. The generated
spectrogram is converted to a waveform by the included vocoder.
Clone the model repository and install its dependencies:
git clone https://huggingface.co/dinhhung1508/SwiftAudio
cd SwiftAudio
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt
python inference.py \
--prompt "Ocean waves with seabirds" \
--seed 42 \
--output ocean.wav
inference.py accepts the following options:
| Option | Description | Default |
|---|---|---|
--prompt | Text description of the requested audio | Required |
--output | Output WAV file path | output.wav |
--seed | Random seed for reproducible generation | 42 |
--model | Hugging Face model ID or local model path | dinhhung1508/SwiftAudio |
--device | Inference device: cuda or cpu | Automatically detected |
python app.py
Open the local URL printed by Gradio, usually
http://127.0.0.1:7860.
import soundfile as sf
from inference import SAMPLE_RATE, generate_audio, load_model
pipeline, vocoder, device, dtype = load_model("dinhhung1508/SwiftAudio")
audio = generate_audio(
pipeline,
vocoder,
prompt="Rain and thunder",
seed=42,
device=device,
dtype=dtype,
)
sf.write("rain.wav", audio, SAMPLE_RATE)
CPU inference is supported but is significantly slower and requires more system memory:
python inference.py \
--device cpu \
--prompt "A train whistles" \
--output train.wav
(1, 4, 32, 128).The checkpoint uses Diffusers' StableDiffusionPipeline as a component
container, but it generates spectrograms rather than natural images. Loading it
as a standard image-generation pipeline without SwiftAudio's post-processing
and vocoder will not produce audio.
SwiftAudio/
├── app.py # Local Gradio demo
├── inference.py # CLI and Python inference API
├── requirements.txt
├── auffusion/ # Vocoder and spectrogram utilities
├── unet/ # One-step student UNet
├── vae/ # Spectrogram VAE
├── vocoder/ # Spectrogram-to-waveform vocoder
├── text_encoder/ and tokenizer/ # CLIP text conditioning
└── scheduler/ # Diffusion scheduler configuration
CUDA out of memory
Close other GPU workloads or use a GPU with at least approximately 10 GB of
available memory. CPU inference can be selected with --device cpu.
The first run takes longer
The first run downloads approximately 5 GB of model files and initializes all pipeline components. Later runs reuse the Hugging Face cache.
The output is an image instead of audio
Use the provided inference.py or app.py. A bare
StableDiffusionPipeline(...) call does not run the included vocoder.
Reproducibility
Use the same prompt, seed, device type, and dependency versions. Results can vary slightly across hardware and library versions.
Rain and thunderOcean waves with seabirdsA train whistlesDishes clattering in a kitchenAdult female is speaking and a young child is cryingSwiftAudio is intended for research and demonstration of fast text-conditioned sound generation, including sound effects, environmental audio, and acoustic scenes.
Users are responsible for evaluating generated audio and ensuring that their use complies with applicable laws, policies, and rights.
The paper is currently under double-blind review:
@article{anonymous2026swiftaudio,
title={SwiftAudio: One-step Text-to-Audio Diffusion-Based Generation
with an Audio-Free Distillation Technique},
author={Anonymous Authors},
journal={Under Review},
year={2026}
}
SwiftAudio uses an Auffusion-compatible latent audio architecture and builds on the Diffusers and Transformers ecosystems.