Downloads · 30 days
0
iconbench/melvo-ace-space
melvo-ace-space is a machine learning model from iconbench. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Downloads · 30 days
0
Access
Public
Updated Oct 11, 2025
Repo size
2.2 MB
Likes
0
Public
Click a slice to open those files.
.png2.2 MB · 70%
From the Hugging Face model README
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
<h1 align="center">✨ ACE-Step ✨</h1> <h1 align="center">🎵 A Step Towards Music Generation Foundation Model 🎵</h1> <p align="center"> <a href="https://ace-step.github.io/">Project</a> | <a href="https://github.com/ace-step/ACE-Step">Code</a> | <a href="https://huggingface.co/ACE-Step/ACE-Step-v1-3.5B">Checkpoints</a> | <a href="https://huggingface.co/spaces/ACE-Step/ACE-Step">Space Demo</a> </p>We introduce ACE-Step, a novel open-source foundation model for music generation that overcomes key limitations of existing approaches and achieves state-of-the-art performance through a holistic architectural design. Current methods face inherent trade-offs between generation speed, musical coherence, and controllability. For instance, LLM-based models (e.g., Yue, SongGen) excel at lyric alignment but suffer from slow inference and structural artifacts. Diffusion models (e.g., DiffRhythm), on the other hand, enable faster synthesis but often lack long-range structural coherence.
ACE-Step bridges this gap by integrating diffusion-based generation with Sana’s Deep Compression AutoEncoder (DCAE) and a lightweight linear transformer. It further leverages MERT and m-hubert to align semantic representations (REPA) during training, enabling rapid convergence. As a result, our model synthesizes up to 4 minutes of music in just 20 seconds on an A100 GPU—15× faster than LLM-based baselines—while achieving superior musical coherence and lyric alignment across melody, harmony, and rhythm metrics. Moreover, ACE-Step preserves fine-grained acoustic details, enabling advanced control mechanisms such as voice cloning, lyric editing, remixing, and track generation (e.g., lyric2vocal, singing2accompaniment).
Rather than building yet another end-to-end text-to-music pipeline, our vision is to establish a foundation model for music AI: a fast, general-purpose, efficient yet flexible architecture that makes it easy to train sub-tasks on top of it. This paves the way for developing powerful tools that seamlessly integrate into the creative workflows of music artists, producers, and content creators. In short, we aim to build the Stable Diffusion moment for music.
conda create -n ace_step python==3.10
conda activate ace_step
pip install -r requirements.txt
conda install ffmpeg
We've tested ACE-Step on various hardware configurations with the following throughput results:
| Device | 27 Steps | 60 Steps |
|---|---|---|
| NVIDIA A100 | 0.036675 | 0.0815 |
| MacBook M2 Max | 0.44 | |
| NVIDIA RTX 4090 | 0.029 | 0.064 |
seconds cost per generated audio (seconds/audio) For example, to generate a 180-second song, multiply 180 by the seconds cost per generated audio (seconds/audio) for the desired device and step count. This will give you the total time required for the generation process.
![]()
python app.py
python app.py --checkpoint_path /path/to/checkpoint --port 7865 --device_id 0 --share --bf16
--checkpoint_path: Path to the model checkpoint (default: downloads automatically)--port: Port to run the Gradio server on (default: 7865)--device_id: GPU device ID to use (default: 0)--share: Enable Gradio sharing link (default: False)--bf16: Use bfloat16 precision for faster inference (default: True)The ACE-Step interface provides several tabs for different music generation and editing tasks:
📋 Input Fields:
⚙️ Settings:
🚀 Generation: Click "Generate" to create music based on your inputs
The examples/input_params directory contains sample input parameters that can be used as references for generating music.
This project is licensed under Apache License 2.0
ACE-Step enables original music generation across diverse genres, with applications in creative production, education, and entertainment. While designed to support positive and artistic use cases, we acknowledge potential risks such as unintentional copyright infringement due to stylistic similarity, inappropriate blending of cultural elements, and misuse for generating harmful content. To ensure responsible use, we encourage users to verify the originality of generated works, clearly disclose AI involvement, and obtain appropriate permissions when adapting protected styles or materials. By using ACE-Step, you agree to uphold these principles and respect artistic integrity, cultural diversity, and legal compliance. The authors are not responsible for any misuse of the model, including but not limited to copyright violations, cultural insensitivity, or the generation of harmful content.
This project is co-led by ACE Studio and StepFun.
If you find this project useful for your research, please consider citing:
@misc{gong2025acestep,
title={ACE-Step: A Step Towards Music Generation Foundation Model},
author={Junmin Gong, Wenxiao Zhao, Sen Wang, Shengyuan Xu, Jing Guo},
howpublished={\url{https://github.com/ace-step/ACE-Step}},
year={2025},
note={GitHub repository}
}