Downloads · 30 days
13
1% of all-time downloads
lichang0928/QA-MDT
QA-MDT is a text-to-audio model from lichang0928. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as mit.
This model, QA-MDT, allows for easy setup and usage for generating music from text prompts. It incorporates a quality-aware training strategy to improve the fidelity of generated music.
Downloads · 30 days
13
1% of all-time downloads
All-time downloads
1.5K
Public
Repo size
14.7 GB
Likes
14
Public
Click a slice to open those files.
.ckpt13.5 GB · 100%
From the Hugging Face model README
This model, QA-MDT, allows for easy setup and usage for generating music from text prompts. It incorporates a quality-aware training strategy to improve the fidelity of generated music.
A Hugging Face Diffusers implementation is available at this model and this space. For more detailed instructions and the official PyTorch implementation, please refer to the project's Github repository and project page.
The model was presented in the paper QA-MDT: Quality-aware Masked Diffusion Transformer for Enhanced Music Generation.